arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 544 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 544 篇

2510.24870 2026-06-02 cs.CL cs.CV cs.IR 89%

Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation

看穿MiRAGE:评估多模态检索增强生成

Alexander Martin, William Walden, Reno Kriz, Dengjia Zhang, Kate Sanders, Eugene Yang, Chihsheng Jin, Benjamin Van Durme

机构 * Johns Hopkins University(约翰霍普金斯大学) Human Language Technology Center of Excellence(人类语言技术卓越中心)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval augmented generation(title);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

AI总结 提出MiRAGE框架,通过InfoF1和CiteF1指标评估多模态RAG的事实性和引用支持,并验证其与人工判断的一致性。

Comments https://github.com/alexmartin1722/mirage

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22829 2026-05-25 cs.IR cs.AI 89%

LFRAG: Layout-oriented Fine-grained Retrieval-Augmented Generation on Multimodal Document Understanding

LFRAG:面向布局的多模态文档理解中的细粒度检索增强生成

Yifan Zhu, Yu Mi, Yue Lu, Yanchu Guan, Zhixuan Chu

机构 * Zhejiang University(浙江大学) Hangzhou High-Tech Zone (Binjiang) Zhejiang University Institute of Blockchain and Data Security(杭州高新技术区(滨江)浙江大学区块链与数据安全研究院)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(title,abstract);分类 cs.IR、cs.AI

AI总结 提出LFRAG框架,通过布局分割和语义-布局融合编码器实现从页面级到块级的细粒度检索,提升多模态文档检索精度并减少生成冗余,在LFDocQA基准上取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10798 2026-07-14 cs.CL 新提交 89%

Trust Before Fusion: QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG

融合前的信任:用于受污染多模态RAG的QIMG-7和源感知分辨率

Saadeldine Eletter, Owais Aijaz, Preslav Nakov

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.CL

AI总结 研究多模态检索增强生成中受污染内容问题,提出QIMG-7基准。针对朴素多模态融合脆弱问题,提出源感知信任分辨率(SATR),Field-Selector变体效果最佳,结果支持选择性信任,显式文本可靠性建模是收益主要驱动因素。

Comments 23 pages, 6 figures, 23 tables. Preprint under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29956 2026-05-29 cs.IR 89%

Uncertainty Quantification for Multimodal Retrieval Augmented Generation

多模态检索增强生成的不确定性量化

Simon Binz, Heydar Soudani, Faegheh Hasibi

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval augmented generation(title,abstract);分类 cs.IR

AI总结 提出 LeMUQ 方法,通过多模态和检索感知的概率信号建模不确定性,提升多模态 RAG 系统的可靠性,AUROC 平均提升 3.8%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10253 2026-05-12 cs.CR cs.AI 89%

Knowledge Poisoning Attacks on Medical Multi-Modal Retrieval-Augmented Generation

针对医疗多模态检索增强生成的知识污染攻击

Peiru Yang, Haoran Zheng, Tong Ju, Shiting Wang, Wanchun Ni, Jiajun Liu, Shangguang Wang, Yongfeng Huang, Tao Qi

机构 * Tsinghua University(清华大学) Beijing University of Posts and Telecommunications(北京邮电大学) Northwestern Polytechnical University(西北工业大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(title,abstract);分类 cs.AI

AI总结 本文提出M³Att框架,针对医疗多模态RAG系统,通过在文本数据中注入隐蔽虚假信息并利用配对视觉数据作为查询无关触发器,提升检索概率,从而在不依赖用户查询先验知识的情况下实现知识污染攻击。

Comments Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20821 2026-04-13 cs.CL cs.AI cs.CE 88%

MultiFinRAG: An Optimized Multimodal Retrieval-Augmented Generation (RAG) Framework for Financial Question Answering

MultiFinRAG:一种优化的多模态检索增强生成(RAG)框架,用于金融问答

Chinmay Gondhalekar, Urjitkumar Patel, Fang-Chun Yeh

机构 * S&P Global Ratings(标普全球评级)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(title,abstract);分类 cs.CL、cs.AI

AI总结 MultiFinRAG通过多模态提取和分层回退策略,提升金融问答任务中跨模态推理的准确性,比ChatGPT-4o在复杂任务上高19个百分点。

Comments Preprint Copy

Journal ref 2025 IEEE International Conference on Big Data (BigData), Macau, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12330 2025-04-18 cs.CL cs.AI 88%

HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation

Pei Liu, Xin Liu, Ruoyu Yao, Junming Liu, Siyuan Meng, Ding Wang, Jun Ma

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(title);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26916 2026-06-26 cs.CV 新提交 88%

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation

PhysRAG: 通过检索增强生成提升视频生成中的物理感知

Kexu Cheng, Zicheng Liu, Mingju Gao, Chunhe Song, Hao Tang

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) Institute of AI for Industries, Chinese Academy of Sciences(中国科学院人工智能产业研究院)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(title,abstract)

AI总结 提出PhysRAG管道,利用检索增强生成(RAG)和可学习查询注入物理知识,在7K高质量视频上训练,在PhyGenBench和VBench上达到最优视觉质量和物理规则遵从性。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17394 2026-04-21 cs.CV 88%

LVLM-Aware Multimodal Retrieval for RAG-Based Medical Diagnosis with General-Purpose Models

面向RAG的医疗诊断的LVLM感知多模态检索

Nir Mazor, Tom Hope

机构 * School of Computer Science and Engineering, The Hebrew University of Jerusalem(希伯来大学耶路撒冷分校计算机科学与工程学院) The Allen Institute for AI (AI2)(人工智能研究院(AI2))

专题命中 多模态RAG :RAG(title,title_cn);retriever(abstract)

AI总结 本文提出轻量级机制提升检索增强LVLM的诊断性能,通过轻量微调和通用模型实现临床分类和VQA任务的竞争力结果,并分析不一致检索预测问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15190 2026-02-18 cs.CL 88%

AIC CTU@AVerImaTeC: dual-retriever RAG for image-text fact checking

AIC CTU@AVerImaTeC:双检索器RAG用于图像-文本事实核查

Herbert Ullrich, Jan Drchal

专题命中 多模态RAG :RAG(title,abstract);retriever(title);retrieval-augmented generation(abstract);分类 cs.CL

AI总结 AIC CTU@AVerImaTeC提出双检索器RAG方法,通过结合文本和图像检索模块,实现高效的图像-文本事实核查,具有低运行成本和易复现性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01513 2026-01-08 cs.CV cs.AI 88%

FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented Generation

FastV-RAG: 向快速且细粒度的视频问答迈进:基于检索增强生成的框架

Gen Li, Peiyu Liu

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) University of International Business and Economics(国际商务经济大学)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(title,abstract);分类 cs.AI

AI总结 FastV-RAG通过引入推测解码和基于相似性的过滤策略,提升视频问答任务的效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23990 2025-11-13 cs.AI 88%

Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding

Mingyang Mao, Mariela M. Perez-Cabarcas, Utteja Kallakuri, Nicholas R. Waytowich, Xiaomin Lin, Tinoosh Mohsenin

机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16636 2025-08-18 cs.CL cs.CV 88%

Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

Yin Wu, Quanyu Long, Jing Li, Jianfei Yu, Wenya Wang

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(title);retrieval-augmented generation(abstract);分类 cs.CL

Comments 21 pages, 6 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14011 2025-04-22 cs.CV cs.AI cs.MM 88%

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation

Fulvio Sanguigni, Davide Morelli, Marcella Cornia, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Pisa(比萨大学)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(title,abstract);分类 cs.AI

Comments IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03995 2025-01-08 cs.LG cs.CV cs.IR cs.IT math.IT 88%

RAG-Check: Evaluating Multimodal Retrieval Augmented Generation Performance

Matin Mortaheb, Mohammad A. Amir Khojastepour, Srimat T. Chakradhar, Sennur Ulukus

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(title);retrieval-augmented generation(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04231 2026-06-04 cs.CL cs.AI 88%

MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A

MM-BizRAG:面向通用企业问答的多模态检索增强生成再思考

Hanoz Bhathena, Parin Rajesh Jhaveri, Rohan Mittal, Prateek Singh, Aymen Kallala, Rachneet Kaur, Yiqiao Jin, Zhen Zeng, Adwait Ratnaparkhi, Denis Kochedykov

机构 * JPMorgan Chase & Co.(摩根大通公司) Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract,abstract_cn);retriever(abstract);分类 cs.CL、cs.AI

AI总结 提出MM-BizRAG框架,通过文档结构感知分割和布局感知解析,结合统一LLM驱动的工件转换与推理时多模态组装,无需微调即可提升企业文档问答性能,在异构企业数据集和两个公开基准上超越基线最多32个百分点。

Comments Accepted at ACL 2026 (Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04502 2025-09-08 cs.CL cs.AI 88%

VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples

Qixin Sun, Ziqin Wang, Hengyuan Zhao, Yilin Li, Kaiyou Song, Linjiang Huang, Xiaolin Hu, Qingpei Guo, Si Liu

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(abstract);retrieval-augmented generation(abstract);retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22868 2026-08-04 cs.CV 版本更新 88%

Seeing the Unseen: Towards Training-Free Inspection for Wind Turbine Blades Using Knowledge-Augmented Vision Language Models

看见不可见:基于知识增强视觉语言模型的无训练风机叶片检测方法

Yang Zhang, Qianyu Zhou, Farhad Imani, Jiong Tang

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract,abstract_cn);retriever(abstract)

AI总结 该研究提出结合RAG与VLM的零样本风机叶片检测框架,构建多模态知识库,在小样本测试中表现优于基线,为工业检测提供数据高效方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02047 2026-03-03 cs.CV 88%

NICO-RAG: Multimodal Hypergraph Retrieval-Augmented Generation for Understanding the Nicotine Public Health Crisis

NICO-RAG:多模态超图检索增强生成用于理解尼古丁公共卫生危机

Manuel Serna-Aguilera, Raegan Anderes, Page Dobbs, Khoa Luu

机构 * University of Arkansas(亚拉巴马大学) University of Arkansas for Medical Sciences(亚拉巴马大学医学科学分校)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(title,abstract)

AI总结 NICO-RAG通过多模态超图检索增强生成,解决尼古丁公共卫生危机中的信息检索与生成问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05141 2025-06-09 eess.AS cs.SD 88%

Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation

Mu Yang, Bowen Shi, Matthew Le, Wei-Ning Hsu, Andros Tjandra

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(title,abstract)

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21968 2026-06-23 cs.CV cs.CL 新提交 87%

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG

先看再缩放:视觉RAG中分辨率-上下文权衡的自适应路由

Oanh N. Tran, Thanh Quoc Hung Le, Oscar Chew, Kuan-Hao Huang, Khoa D. Doan

机构 * VinUni-Illinois Smart Health Center, VinUniversity(VinUni-Illinois智慧健康中心,VinUniversity) Texas A&M University(德克萨斯农工大学)

专题命中 多模态RAG :RAG(title,title_cn);分类 cs.CL

AI总结 提出ViRGo框架,通过自适应路由选择全局感知、基于补丁或注意力的检索,以平衡小目标细节和大目标上下文,提升视觉RAG的精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12323 2025-10-15 cs.AI 87%

RAG-Anything: All-in-One RAG Framework

Zirui Guo, Xubin Ren, Lingrui Xu, Jiahao Zhang, Chao Huang

机构 * The University of Hong Kong(香港大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);knowledge retrieval(abstract);hybrid retrieval(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24875 2026-07-29 cs.LG 新提交 87%

FinAbstain: Uncertainty-Calibrated Multimodal RAG for Selective Financial Forecasting

FinAbstain:用于选择性金融预测的不确定性校准多模态检索增强生成模型

Dorothy Torres, Wei Cheng, Henan Huang

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);retriever(abstract)

AI总结 研究针对大语言模型在金融预测中证据不足时高置信度的问题,提出FinAbstain框架,通过多模态检索增强生成及选择性预测,经多种不确定性评估方法和指标评估,贡献了时间安全架构、复合不确定性公式和可重复评估蓝图。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04625 2026-07-07 cs.CV cs.AI 新提交 86%

Hierarchical Evidence-Driven Reasoning for Long Document Understanding

用于长文档理解的分层证据驱动推理

Junyu Xiong, Yonghui Wang, Rongjian Gu, Chenyu Liu, Bing Yin, Wengang Zhou, Houqiang Li

机构 * University of Science and Technology of China(中国科学技术大学) iFLYTEK Research(科大讯飞研究院)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.AI

AI总结 针对现有多模态RAG管道面临的问题,提出分层证据驱动的多模态RAG框架HIEVI-RAG,通过四阶段管道,包括分层问题分解、粗视觉页面检索、细粒度页面验证和内存引导迭代生成,提升长文档理解,效果显著优于现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22217 2026-02-27 cs.IR cs.AI 86%

RAGdb: A Zero-Dependency, Embeddable Architecture for Multimodal Retrieval-Augmented Generation on the Edge

RAGdb: 一种无依赖、可嵌入的多模态检索增强生成边缘计算架构

Ahmed Bin Khalid

机构 * SKAS IT

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);vector search(abstract);分类 cs.IR、cs.AI

AI总结 RAGdb通过单文件SQLite容器实现多模态检索增强生成,无需GPU推理,提升边缘计算效率与数据主权

Comments 6 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15435 2025-11-20 cs.CV cs.AI cs.IR 86%

HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation

HV-Attack:多模态检索增强生成的分层视觉攻击

Linyin Luo, Yujuan Ding, Yunshan Ma, Wenqi Fan, Hanjiang Lai

机构 * The Hong Kong Polytechnic University(香港理工大学) Sun Yat-Sen University(中山大学) Singapore Management University(新加坡管理学院)

专题命中 多模态RAG :retrieval augmented generation(title);retrieval-augmented generation(abstract);RAG(abstract);retriever(abstract)

AI总结 本文提出了一种分层视觉攻击方法,通过在图像输入中添加不可察觉扰动,破坏多模态检索增强生成系统的检索和生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15253 2026-04-21 cs.CL cs.CV 86%

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding

超越上下文的规模:文档理解的多模态检索增强生成综述

Sensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan, Yong Xien Chng, Qing-Guo Chen, Weihua Luo, Kaifu Zhang, Jia-Wang Bian, Mingming Gong

机构 * MBZUAI Alibaba Group(阿里巴巴集团) Tsinghua University(清华大学) Wuhan University(武汉大学) Nanyang Technological University(南洋理工大学) University of Melbourne(墨尔本大学)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract,abstract_cn);分类 cs.CL

AI总结 本文综述了多模态检索增强生成在文档理解中的应用,提出基于领域、检索模态和粒度的分类体系,总结了关键数据集、基准测试和行业应用,指出了效率、细粒度表示和鲁棒性等挑战。

Comments Accepted by ACL2026 Main Conference; Project is available at https://github.com/SensenGao/Multimodal-RAG-Survey-For-Document

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27259 2026-06-23 cs.CV 版本更新 86%

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

看到场景才重要:通过场景感知的长视频基准揭示视频理解模型中的遗忘现象

Seng Nam Chen, Hao Chen, Chenglam Ho, Xinyu Mao, Jinping Wang, Yu Zhang, Chao Li

机构 * CUHK (SZ)(香港中文大学(深圳)) University of Cambridge(剑桥大学) UESTC(电子科技大学) CUHK(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract,abstract_cn)

AI总结 本文提出SceneBench基准,揭示视频理解模型在长场景上下文中的遗忘问题,并提出Scene-RAG方法提升性能2.50%。

Comments Accepted to CVPR 2026 (Highlight)

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21544 2026-08-03 cs.CV cs.CL 85%

Vision Meets Language: A RAG-Augmented YOLOv8 Framework for Coffee Disease Diagnosis and Farmer Assistance

Semanto Mondal

机构 * University of Naples Federico II(那不勒斯费德里科二世大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(abstract);retrieval-augmented generation(abstract);分类 cs.CL

Comments There are 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01737 2026-06-02 cs.AI 85%

TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination

TrafficRAG:用于交通事故责任认定的多模态RAG框架

Xu Li, Zedong Fu, Xinyi Li, Xun Han

机构 * Southwest Petroleum University(西南石油大学) Sichuan Police College(四川警察学院)

专题命中 多模态RAG :RAG(title,title_cn);hybrid retrieval(abstract);分类 cs.AI

AI总结 提出TrafficRAG框架,通过视觉语言模型生成结构化描述、混合检索获取法规和案例、大语言模型融合多模态证据进行推理,实现自动化交通事故责任分析报告生成。

Comments 12 pages, 3 figures, accepted at ICANN 2026

详情

展开后加载摘要…

URL PDF HTML 收藏