Comments8 pages, 10 figures, 4 tables. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026). Project page and source code: https://github.com/LZY-1021/RoboHarness
CommentsAdded a conceptual diagram for the LGC architecture, 14 pages, 10 figures, 7 tables. Submitted to IEEE Transactions on Information Forensics and Security. The source code is available at https://github.com/eihmuekhine/Latent-Geometric-Chords
VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
VisRAG2.0:通过视觉检索增强生成中的证据引导多图像推理减轻视觉幻觉
Yubo Sun, Chunyi Peng, Yukun Yan, Shi Yu, Zhenghao Liu, Sen Mei, Chi Chen, Maosong Sun
机构
*
School of Software and Microelectronics, Peking University, China(北京大学软件与微电子学院)
;
School of Computer Science and Engineering, Northeastern University, China(东北大学计算机科学与工程学院)
;
Department of Computer Science and Technology, Institute for AI, Tsinghua University, China(清华大学人工智能研究院计算机科学与技术系)
机构
*
School of Computer Science, Peking University(北京大学计算机科学学院)
;
School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院)
;
School of Information, Renmin University of China(中国人民大学信息学院)
;
School of Integrated Circuit Science and Engineering, Beihang University(北京航空航天大学集成电路科学与工程学院)
CommentsAccepted at the 42nd IEEE International Conference on Software Maintenance and Evolution (ICSME 2026), Tool Demonstration and Data Showcase Track
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
GUIDE:通过实时网络视频检索和即插即用标注解决GUI代理的领域偏见
Rui Xie, Zhi Gao, Chenrui Shi, Zirui Shang, Lu Chen, Qing Li
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
State Key Laboratory for General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,北京通用人工智能研究院)
;
Beijing Institute of Technology(北京理工大学)
CommentsAccepted to ECCV 2026. 30 pages: 15-page main paper followed by supplementary material as an appendix (Sections A-F). Project page: https://sharryXR.github.io/GUIDE/
PRIMA: Pre-Training with Risk-Integrated Image--Metadata Alignment for Medical Diagnosis with LLM-Based Feature Aggregation
PRIMA:通过大语言模型进行风险集成图像-元数据对齐的医学诊断预训练
Yiqing Wang, Chunming He, Ziyun Yang, Maria Woodward, Ming-Chen Lu, Mercy Pawar, Leslie Niziol, Sina Farsiu
机构
*
Department of Biomedical Engineering, Duke University(杜克大学生物医学工程系)
;
Department of Ophthalmology and Visual Sciences, University of Michigan(密歇根大学眼科与视觉科学系)
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
Wiki-R1: 通过数据和采样课程激励基于知识的多模态推理用于VQA
Shan Ning, Longtian Qiu, Xuming He
机构
*
ShanghaiTech University(上海科技大学)
;
Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程研究中心)
;
Lingang Laboratory(临港实验室)