arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 544 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 544 篇

2605.05594 2026-05-08 cs.CL cs.CV cs.LG 85%

The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation

上下文的成本:减轻多模态检索增强生成中的文本偏差

Hoin Jung, Xiaoqian Wang

机构 * Elmore Family School of Electrical and Computer Engineering(埃尔莫夫家族电气与计算机工程学院)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract,abstract_cn);分类 cs.CL

AI总结 本文探讨了多模态大语言模型在整合检索增强生成时出现的文本偏差问题,提出BAIR方法通过恢复视觉显著性并施加位置感知惩罚来提升多模态接地和诊断可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27724 2026-05-01 cs.AI 85%

Iterative Multimodal Retrieval-Augmented Generation for Medical Question Answering

迭代多模态检索增强生成用于医疗问答

Xupeng Chen, Binbin Shi, Chenqian Le, Jiaqi Zhang, Kewen Wang, Ran Gong, Jinhan Zhang, Chihang Wang

机构 * New York University, New York, USA(纽约大学) Tsinghua University, Beijing, China(清华大学) Independent Researcher(独立研究者) Chinese Academy of Sciences, Beijing, China(中国科学院)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract,abstract_cn);分类 cs.AI

AI总结 MED-VRAG通过多模态检索增强生成框架,利用文档页面图像而非OCR文本进行医疗问答,提升准确率至78.6%,并验证检索和迭代对问答性能的积极影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12735 2026-04-22 cs.CV cs.CL 85%

VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph

VimRAG:通过多模态记忆图导航大规模视觉上下文

Qiuchen Wang, Shihang Wang, Yu Zeng, Qiang Zhang, Fanrui Zhang, Zhuoning Guo, Bosi Zhang, Wenxuan Huang, Lin Chen, Zehui Chen, Pengjun Xie, Ruixue Ding

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract,abstract_cn);分类 cs.CL

AI总结 VimRAG通过多模态记忆图提升检索增强生成在处理大规模视觉上下文中的能力,采用动态有向无环图结构和图调制视觉记忆编码机制,实现高效的信息检索与推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10030 2026-03-03 cs.CR cs.AI 85%

Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment

在RAG-as-a-Service环境中保护多模态知识版权

Tianyu Chen, Jian Lou, Wenjie Wang

机构 * ShanghaiTech University(上海科技大学) Sun Yat-sen University(中山大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.AI

AI总结 AQUA是一种用于多模态RAG系统中图像知识保护的首个水印框架,通过嵌入语义信号实现高效、隐蔽的版权追溯。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09721 2025-09-15 cs.CV cs.AI cs.LG 85%

A Multimodal RAG Framework for Housing Damage Assessment: Collaborative Optimization of Image Encoding and Policy Vector Retrieval

Jiayi Miao, Dingxin Lu, Zhuqi Wang

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08826 2025-06-03 cs.CL cs.AI cs.IR 85%

Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation

Mohammad Mahdi Abootorabi, Amirhosein Zobeiri, Mahdi Dehghani, Mohammadali Mohammadkhani, Bardia Mohammadi, Omid Ghahroodi, Mahdieh Soleymani Baghshah, Ehsaneddin Asgari

机构 * Computer Engineering Department, Sharif University of Technology(沙斐大学计算机工程系) College of Interdisciplinary Science and Technology, University of Tehran(德黑兰大学跨学科科学与技术学院) Computer Engineering Department, K.N. Toosi University of Technology(技术学院计算机工程系) Qatar Computing Research Institute(卡塔尔计算研究院)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract,comments);分类 cs.IR、cs.CL、cs.AI

Comments GitHub repository: https://github.com/llm-lab-org/Multimodal-RAG-Survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22833 2026-05-25 cs.IR cs.AI cs.LG 85%

RAG4Outcome: A Retrieval-Augmented Multimodal Framework for Prognostic Prediction in Chronic Osteomyelitis

RAG4Outcome:用于慢性骨髓炎预后预测的检索增强多模态框架

Daqian Shi, Pei Han, Jishizhan Chen, Yang Wang, Xiaolei Diao, Xianyou Zheng, Pengfei Cheng

机构 * Queen Mary University of London(女王玛丽大学) Shanghai Sixth People’s Hospital Affiliated to SJTU School of Medicine(上海第六人民医院附属复旦大学医学院) University College London(大学学院伦敦)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 提出RAG4Outcome框架,通过检索增强生成技术融合PET-CT影像报告、结构化手术诊断记录和非结构化随访笔记等多模态临床数据,实现慢性骨髓炎的预后预测,提高可解释性和临床可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23276 2026-04-28 cs.CV cs.AI cs.CL 85%

Lightweight and Production-Ready PDF Visual Element Parsing

轻量且适用于生产的PDF视觉元素解析

Meizhu Liu, Yassi Abbasi, Matthew Rowe, Michael Avendi, Paul Li

机构 * Oracle AI

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种轻量级PDF解析框架,通过空间启发式、布局分析和语义相似性结合,实现高准确率的视觉元素检测与标题关联,提升多模态RAG检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03354 2026-06-03 cs.CR 85%

ImageAuditor: Membership Inference Attack against Image-based Retrieval-Augmented Generation

ImageAuditor: 针对基于图像的检索增强生成系统的成员推理攻击

Jinghuai Zhang, Pengyue Yu, Zhexiao Lin, Kunlin Cai, Fnu Suya, Yuan Tian

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract,abstract_cn)

AI总结 提出首个针对基于图像的检索增强生成(IRAG)的成员推理攻击方法ImageAuditor,通过奖励引导策略优化解决跨模态检索和判别信号提取两大挑战,在仅需四次查询时AUROC超过80%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05285 2025-07-15 cs.CL cs.AI cs.CY cs.IR 85%

Beyond classical and contemporary models: a transformative AI framework for student dropout prediction in distance learning using RAG, Prompt engineering, and Cross-modal fusion

Miloud Mihoubi, Meriem Zerkouk, Belkacem Chikhaoui

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

Comments 13 pages, 8 figures, 1 Algorithms, 17th International Conference on Education and New Learning Technologies,: 30 June-2 July, 2025 Location: Palma, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02544 2025-06-06 cs.CL cs.AI cs.IR 85%

CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG

Yang Tian, Fan Liu, Jingyuan Zhang, Victoria W., Yupeng Hu, Liqiang Nie

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

Comments Accepted to ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05874 2025-05-30 cs.CV cs.AI cs.CL cs.IR cs.LG 85%

VideoRAG: Retrieval-Augmented Generation over Video Corpus

Soyeong Jeong, Kangsan Kim, Jinheon Baek, Sung Ju Hwang

机构 * KAIST(韩国科学技术院)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.IR、cs.CL、cs.AI

Comments ACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16086 2025-04-30 cs.IR cs.AI cs.CL cs.CV eess.IV 85%

Towards Interpretable Radiology Report Generation via Concept Bottlenecks using a Multi-Agentic RAG

Hasan Md Tusfiqur Alam, Devansh Srivastav, Md Abdul Kadir, Daniel Sonntag

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) University of Oldenburg(奥尔登堡大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

Comments Accepted in the 47th European Conference for Information Retrieval (ECIR) 2025

Journal ref Lecture Notes in Computer Science (LNCS) 2025, Volume 15574

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08748 2025-04-15 cs.IR cs.AI cs.CL cs.ET cs.LG 85%

A Survey of Multimodal Retrieval-Augmented Generation

Lang Mei, Siyu Mo, Zhihan Yang, Chong Chen

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.IR、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15262 2024-12-23 cs.CL cs.AI cs.IR 85%

Advanced ingestion process powered by LLM parsing for RAG system

Arnau Perez, Xavier Vizcaino

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08935 2026-08-11 cs.AI 新提交 84%

Integrated Multimodal AI System for Retrieval-Augmented Reasoning, Object Sensing, and Damage Analysis

用于检索增强推理、物体感知与损伤分析的集成多模态AI系统

Kalelo Dukuray, Israel Pina, Evan Perez, Erika Ardiles-Cruz, Jie Wei

机构 * City College of New York(纽约城市学院) Air Force Research Lab(空军研究实验室)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本研究提出整合RAG、热谱感知等技术的多模态AI系统,用于损伤评估,通过多模块对比实验验证了其在提升推理准确性、鲁棒性及跨场景检测方面的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08883 2026-08-11 cs.AI 新提交 84%

AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups

AquiLLM:支持研究群体中隐性知识捕获的架构

Jack Stark, Srinath Saikrishnan, Vikram Seenivasan, Bernie Boscoe, Andrew Lizarraga, Tuan Do

机构 * Southern Oregon University(南俄勒冈大学) University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 AquiLLM是一款采用开放权重模型的开源模块化RAG-LLM框架,经领域专家反馈优化了架构与功能,可支持研究群体捕获隐性知识,助力AI贴合科研实践。

Comments Accepted for publication in the NGEN-AI 2026 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15838 2026-06-16 cs.IR 新提交 84%

Intelligent Multimodal Retrieval and Reasoning for Geospatial Knowledge Discovery on the I-GUIDE Platform

面向I-GUIDE平台地理空间知识发现的智能多模态检索与推理

Yunfan Kang, Erick Li, Furqan Baig, Wei Hu, Alexander Michels, Anand Padmanabhan, Shaowen Wang

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.IR

AI总结 提出I-GUIDE Smart Search系统,结合多模态索引与知识图谱的迭代RAG管道,实现地理空间异构数据的语义检索、图谱溯源和对话合成,在单A100部署下支持约100并发用户,提升检索覆盖与答案质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00360 2026-06-03 cs.CL 84%

CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA

CourseTimeQA: 一个讲座视频基准和一种用于时间戳问答的延迟约束跨模态融合方法

Vsevolod Kovalev, Parteek Kumar

专题命中 多模态RAG :RAG(summary_cn,abstract);retriever(abstract);分类 cs.CL

AI总结 针对教育讲座视频中的时间戳问答任务,在单GPU延迟/内存预算下,提出了CourseTimeQA基准和一种轻量级延迟约束跨模态检索器CrossFusion-RAG,通过冻结编码器、浅层查询无关交叉注意力和时间一致性正则化,在nDCG@10和MRR上分别提升0.10和0.08,中位端到端延迟约1.55秒。

Comments This paper is being withdrawn because an error in our measurement procedure produced incorrect values in our reported retrieval results (Tables I, II, V, and VI, and the corresponding headline figures in the Abstract). Several of our empirical claims depend on these measurements and therefore do not hold as stated

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15663 2026-04-20 cs.SE cs.AI 84%

CodeMMR: Bridging Natural Language, Code, and Image for Unified Retrieval

CodeMMR:连接自然语言、代码和图像以实现统一检索

Jiahui Geng, Qing Li, Fengyu Cai, Fakhri Karray

机构 * MBZUAI Linköping University(林雪平大学) University of Groningen(格罗宁根大学) TU Darmstadt(图宾根大学)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出CodeMMR,通过指令式多模态对齐,将自然语言、代码和图像嵌入共享语义空间,实现跨模态和语言的强泛化,优于基线模型,并提升RAG中的代码生成精度与视觉 grounding。

Journal ref CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06179 2026-04-09 cs.IR cs.CL 84%

ARIA: Adaptive Retrieval Intelligence Assistant -- A Multimodal RAG Framework for Domain-Specific Engineering Education

ARIA:自适应检索智能助手——面向领域特定工程教育的多模态RAG框架

Yue Luo, Dibakar Roy Sarkar, Rachel Herring Sangree, Somdatta Goswami

机构 * Dalian University of Technology(大连理工大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

AI总结 ARIA通过多模态内容提取管道和e5-large-v2模型,实现领域特定工程教育的智能教学助手,展现高精度和教学一致性,验证了其在课程相关问题上的高准确率和响应质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02132 2026-04-03 cs.CL cs.CR cs.CV cs.IR 84%

One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image

一张图片足矣:通过单张图片对视觉文档检索增强生成进行污染攻击

Ezzeldin Shereen, Dan Ristea, Shae McFadden, Burak Hasircioglu, Vasilios Mavroudis, Chris Hicks

机构 * The Alan Turing Institute(艾伦·图灵研究所) University College London(伦敦大学学院)

专题命中 多模态RAG :retrieval augmented generation(title);retrieval-augmented generation(abstract);RAG(abstract);分类 cs.IR、cs.CL

AI总结 本文研究了视觉文档检索增强生成(VD-RAG)对污染攻击的脆弱性,通过单张恶意图片实现针对性和通用性攻击,展示了VD-RAG在目标和通用设置下的漏洞,但在黑盒攻击下表现出一定的鲁棒性。

Comments Published in Transactions on Machine Learning Research (03/2026)

Journal ref Transactions on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05959 2026-03-24 cs.CL cs.AI cs.CV 84%

M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG

M4-RAG:大规模多语言多文化多模态检索增强生成

David Anugraha, Patrick Amadeus Irawan, Anshul Singh, En-Shiun Annie Lee, Genta Indra Winata

机构 * Stanford University(斯坦福大学) MBZUAI Indian Institute of Science(印度科学研究院) Ontario Tech University(安大略技术大学) University of Toronto(多伦多大学) Capital One

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 M4-RAG提出一个覆盖42种语言、56种方言和189个国家的多模态大规模基准,通过构建8万多个文化多样化的图像-问题对,评估跨语言和模态的检索增强视觉问答性能,揭示模型大小与检索效果的不匹配问题。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08181 2025-11-18 cs.IR cs.AI 84%

MARC: Multimodal and Multi-Task Agentic Retrieval-Augmented Generation for Cold-Start Recommender System

Seung Hwan Cho, Yujin Yang, Danik Baeck, Minjoo Kim, Young-Min Kim, Heejung Lee, Sangjin Park

机构 * Department of Industrial Data Engineering, Hanyang University, Republic of Korea(工业数据工程系,翰阳大学) School of Interdisciplinary Industrial Studies, Hanyang University, Republic of Korea(跨学科工业研究学院,翰阳大学)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.IR、cs.AI

Comments 13 pages, 2 figures, Accepted at RDGENAI at CIKM 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15418 2025-11-10 cs.CL cs.AI 84%

Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs

Lee Qi Zun, Mohamad Zulhilmi Bin Abdul Halim, Goh Man Fye

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04145 2025-10-07 cs.CV cs.CL cs.IR 84%

Automating construction safety inspections using a multi-modal vision-language RAG framework

Chenxin Wang, Elyas Asadi Shamsabadi, Zhaohui Chen, Luming Shen, Alireza Ahmadian Fard Fini, Daniel Dias-da-Costa

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

Comments 33 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20769 2025-09-26 cs.IR cs.AI cs.CV 84%

Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems

Tuo Zhang, Yuechun Sun, Ruiliang Liu

机构 * Museus University of Science and Technology of China(中国科学技术大学) British Museum(大英博物馆)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10571 2025-09-23 cs.AI cs.CL 84%

Agentic AI with Orchestrator-Agent Trust: A Modular Visual Classification Framework with Trust-Aware Orchestration and RAG-Based Reasoning

Konstantinos I. Roumeliotis, Ranjan Sapkota, Manoj Karkee, Nikolaos D. Tselikas

机构 * University of the Peloponnese, Department of Informatics and Telecommunications(希腊皮埃蒙特大学信息与电信系) Cornell University, Department of Biological and Environmental Engineering(康奈尔大学生物与环境工程系)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16701 2025-09-03 cs.IR cs.CL 84%

AlzheimerRAG: Multimodal Retrieval Augmented Generation for Clinical Use Cases using PubMed articles

Aritra Kumar Lahiri, Qinmin Vivian Hu

机构 * Department of Computer Science, Toronto Metropolitan University(计算机科学系,多伦多 Metropolitan 大学)

专题命中 多模态RAG :retrieval augmented generation(title);retrieval-augmented generation(abstract);RAG(abstract);分类 cs.IR、cs.CL

Journal ref Machine Learning and Knowledge Extraction. 2025; 7(3):89

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09170 2025-08-14 cs.LG cs.AI cs.CV cs.IR 84%

Multimodal RAG Enhanced Visual Description

Amit Kumar Jaiswal, Haiming Liu, Ingo Frommholz

机构 * Indian Institute of Technology (BHU)(印度理工学院(BHU)) University of Southampton(南安普顿大学) Modul University Vienna(维也纳应用科技大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

Comments Accepted by ACM CIKM 2025. 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏