VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
VisRAG2.0:通过视觉检索增强生成中的证据引导多图像推理减轻视觉幻觉
Yubo Sun, Chunyi Peng, Yukun Yan, Shi Yu, Zhenghao Liu, Sen Mei, Chi Chen, Maosong Sun
机构
*
School of Software and Microelectronics, Peking University, China(北京大学软件与微电子学院)
;
School of Computer Science and Engineering, Northeastern University, China(东北大学计算机科学与工程学院)
;
Department of Computer Science and Technology, Institute for AI, Tsinghua University, China(清华大学人工智能研究院计算机科学与技术系)
PRIMA: Pre-Training with Risk-Integrated Image--Metadata Alignment for Medical Diagnosis with LLM-Based Feature Aggregation
PRIMA:通过大语言模型进行风险集成图像-元数据对齐的医学诊断预训练
Yiqing Wang, Chunming He, Ziyun Yang, Maria Woodward, Ming-Chen Lu, Mercy Pawar, Leslie Niziol, Sina Farsiu
机构
*
Department of Biomedical Engineering, Duke University(杜克大学生物医学工程系)
;
Department of Ophthalmology and Visual Sciences, University of Michigan(密歇根大学眼科与视觉科学系)
CommentsAccepted at the 43rd International Conference on Machine Learning (ICML 2026) Workshop on Efficient Multimodal Question Answering (EMM-QA), Seoul, South Korea. Copyright 2026 by the author(s). (Archival)