arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 20 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 20 篇

2604.04969 2026-07-14 cs.IR cs.AI 版本更新 93%

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation

MG$^2$-RAG:多粒度图用于多模态检索增强生成

Sijun Dai, Qiang Huang, Xiaoxing You, Jun Yu

机构 * School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机学院)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(title,abstract);分类 cs.IR、cs.AI

AI总结 MG$^2$-RAG通过构建层次化多模态知识图谱,融合文本与视觉信息,提升多模态检索与生成性能,实验显示其在四个任务中均取得最佳效果,同时显著提升效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14445 2026-06-23 cs.CL cs.AI cs.HC cs.LG 版本更新 91%

Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning

Tell Me:基于LLM的心理健康助手,集成RAG、合成对话生成与智能体规划

Trishala Jayesh Ahalpara

机构 * Fujitsu Research of America(富士通美国研究)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 提出Tell Me系统,利用大语言模型通过RAG实现个性化对话、合成客户-治疗师对话生成以及智能体规划生成自我护理计划,旨在提供可及的心理支持而非替代专业治疗。

Comments 8 pages, 2 figures, 1 Table. Submitted to the Computation and Language (cs.CL) category. Uses the ACL-style template. Code and demo will be released at: https://github.com/trystine/Tell_Me_Mental_Wellbeing_System

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01613 2026-06-16 cs.IR cs.AI cs.MA 版本更新 89%

TechRAG: Evidence-Gated Multimodal Agentic RAG for Technical Literature Reasoning

TechGraphRAG:面向技术文献推理的智能图增强RAG框架

Kanwar Bharat Singh

机构 * Global Tire Intelligence and Solutions (GTIS)(全球轮胎智能与解决方案(GTIS)) The Goodyear Tire & Rubber Company(固特异轮胎与橡胶公司)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 提出一种13步自主流水线的智能检索增强生成框架,通过证据充分性评分、知识图谱遍历和自校正生成,支持领域特定技术文献推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22868 2026-08-04 cs.CV 版本更新 88%

Seeing the Unseen: Towards Training-Free Inspection for Wind Turbine Blades Using Knowledge-Augmented Vision Language Models

看见不可见:基于知识增强视觉语言模型的无训练风机叶片检测方法

Yang Zhang, Qianyu Zhou, Farhad Imani, Jiong Tang

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract,abstract_cn);retriever(abstract)

AI总结 该研究提出结合RAG与VLM的零样本风机叶片检测框架,构建多模态知识库,在小样本测试中表现优于基线,为工业检测提供数据高效方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27259 2026-06-23 cs.CV 版本更新 86%

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

看到场景才重要:通过场景感知的长视频基准揭示视频理解模型中的遗忘现象

Seng Nam Chen, Hao Chen, Chenglam Ho, Xinyu Mao, Jinping Wang, Yu Zhang, Chao Li

机构 * CUHK (SZ)(香港中文大学(深圳)) University of Cambridge(剑桥大学) UESTC(电子科技大学) CUHK(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract,abstract_cn)

AI总结 本文提出SceneBench基准,揭示视频理解模型在长场景上下文中的遗忘问题,并提出Scene-RAG方法提升性能2.50%。

Comments Accepted to CVPR 2026 (Highlight)

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24060 2026-08-13 cs.RO 版本更新 83%

RoboHarness: A Memory-Augmented Policy Harness for Vision-Language-Action Model Robustness via In-Context Adaptation

SOMA:通过上下文适应提升视觉-语言-动作模型鲁棒性的战略编排与内存增强系统

Zhuoran Li, Zhiyang Li, Kaijun Zhou, Jinyu Gu

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 SOMA通过对比双记忆检索增强生成(RAG)、归因驱动大语言模型(LLM)编排器和可扩展模型上下文协议(MCP)干预,提升视觉-语言-动作模型在分布外任务中的鲁棒性,实验表明其在长周期任务链中提升了89.1%的绝对成功率。

Comments 8 pages, 10 figures, 4 tables. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026). Project page and source code: https://github.com/LZY-1021/RoboHarness

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31219 2026-06-11 cs.CV cs.CR cs.LG 版本更新 80%

Latent Geometric Chords for Query-Efficient Decision-Based Adversarial Attacks

潜在几何和弦:面向查询高效决策型对抗攻击

Ei Hmue Khine, Yao Li, Jiebao Sun, Shengzhu Shi, Zhichang Guo, Boying Wu

专题命中 多模态RAG :RAG(summary_cn,abstract)

AI总结 提出潜在几何和弦(LGC)方法,通过曲率感知的几何搜索在压缩语义流形中导航决策边界,并引入残差对抗生成(RAG)机制以高视觉保真度实现查询高效的决策型黑盒对抗攻击。

Comments Added a conceptual diagram for the LGC architecture, 14 pages, 10 figures, 7 tables. Submitted to IEEE Transactions on Information Forensics and Security. The source code is available at https://github.com/eihmuekhine/Latent-Geometric-Chords

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09733 2026-07-29 cs.CL cs.CV 版本更新 79%

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation

VisRAG2.0:通过视觉检索增强生成中的证据引导多图像推理减轻视觉幻觉

Yubo Sun, Chunyi Peng, Yukun Yan, Shi Yu, Zhenghao Liu, Sen Mei, Chi Chen, Maosong Sun

机构 * School of Software and Microelectronics, Peking University, China(北京大学软件与微电子学院) School of Computer Science and Engineering, Northeastern University, China(东北大学计算机科学与工程学院) Department of Computer Science and Technology, Institute for AI, Tsinghua University, China(清华大学人工智能研究院计算机科学与技术系)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);分类 cs.CL

AI总结 研究针对视觉检索增强生成中VLM存在的视觉幻觉及证据识别问题,提出证据引导的多图像推理框架EVisRAG,引入RS-GRPO改进训练,实验表明该方法能提升性能、减少幻觉,有效提高视觉基础和推理可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04080 2026-06-30 cs.IR 版本更新 77%

Caption Injection for Optimization in Generative Search Engine

生成搜索引擎中的标题注入用于优化

Xiaolu Chen, Jie Bao, Haojie Wu, Zhen Chen, Yong Liao

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR

AI总结 本文提出Caption Injection,一种多模态G-SEO方法,通过提取图像标题并注入文本内容,提升生成搜索中的主观可见性,实验表明其在G-EVAL指标下优于文本-only基线。

Comments 24 pages, 4 figures, ECML PKDD 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08976 2026-06-16 cs.CV cs.DC cs.IR 版本更新 77%

MIRAGE: Runtime Scheduling for Multi-Vector Image Retrieval with Hierarchical Decomposition

MIRAGE:基于层次分解的多向量图像检索运行时调度

Maoliang Li, Ke Li, Yaoyang Liu, Jiayu Chen, Zihao Zheng, Yinjun Wu, Chenchen Liu, Xiang Chen

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院) School of Information, Renmin University of China(中国人民大学信息学院) School of Integrated Circuit Science and Engineering, Beihang University(北京航空航天大学集成电路科学与工程学院)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval augmented generation(abstract);分类 cs.IR

AI总结 提出MIRAGE框架,通过层次化分解和跨层次相似性一致性减少冗余计算,实现多向量图像检索的精度提升和3.5倍计算加速。

Comments Will appear in DAC'2026, camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14830 2026-07-21 cs.SE 版本更新 75%

AI Prototyper: A Figma Plugin for Decomposition-Based GUI Prototyping with LLMs

AI原型制作器:一个用于基于分解的GUI原型制作的Figma插件

Tawatchai Salangsingha, Ashkan Sami, Md Zia Ullah, Iain McGregor

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract)

AI总结 研究针对GUI原型制作耗时问题,提出AI Prototyper插件,通过分解和检索增强生成管道,结合特定技术栈与大语言模型,经人工参与编辑步骤,能多语言输入,在时间和质量维度上表现良好。

Comments Accepted at the 42nd IEEE International Conference on Software Maintenance and Evolution (ICSME 2026), Tool Demonstration and Data Showcase Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26266 2026-07-01 cs.AI cs.CV 版本更新 70%

GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation

GUIDE:通过实时网络视频检索和即插即用标注解决GUI代理的领域偏见

Rui Xie, Zhi Gao, Chenrui Shi, Zirui Shang, Lu Chen, Qing Li

机构 * Shanghai Jiao Tong University(上海交通大学) State Key Laboratory for General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,北京通用人工智能研究院) Beijing Institute of Technology(北京理工大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);分类 cs.AI

AI总结 GUIDE通过实时网络视频检索和即插即用标注框架,解决GUI代理的领域偏见问题,通过视频语义分析和自动化标注流程提升代理对特定应用的操作流程和UI布局的理解,实验表明其在多代理系统和单模型代理中均能提升性能。

Comments Accepted to ECCV 2026. 30 pages: 15-page main paper followed by supplementary material as an appendix (Sections A-F). Project page: https://sharryXR.github.io/GUIDE/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23297 2026-07-29 cs.CV 版本更新 67%

PRIMA: Pre-Training with Risk-Integrated Image--Metadata Alignment for Medical Diagnosis with LLM-Based Feature Aggregation

PRIMA:通过大语言模型进行风险集成图像-元数据对齐的医学诊断预训练

Yiqing Wang, Chunming He, Ziyun Yang, Maria Woodward, Ming-Chen Lu, Mercy Pawar, Leslie Niziol, Sina Farsiu

机构 * Department of Biomedical Engineering, Duke University(杜克大学生物医学工程系) Department of Ophthalmology and Visual Sciences, University of Michigan(密歇根大学眼科与视觉科学系)

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

AI总结 提出PRIMA框架,通过检索增强生成策展风险-疾病关联专家语料库优化文本编码器,用双编码器预训练策略及四个互补损失函数弥合模态差距,融合特征用于疾病分类,性能优于其他方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04683 2026-08-05 cs.CV cs.CL 版本更新 57%

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models

是看不见还是不知道?归因视觉语言模型中的错误

Khang Nhat Hoang Vo, Artem Vazhentsev, Artem Shelmanov, Timothy Baldwin, Yova Kementchedjhieva

机构 * MBZUAI The University of Melbourne(MBZUAI墨尔本大学)

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

AI总结 研究视觉语言模型在回答需额外知识问题时的错误,提出统一框架分离失败模式,探讨预生成信号能否预测错误源,发现可在解码前预测,能据此进行针对性干预。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15782 2026-08-04 cs.AI cs.CV 版本更新 57%

Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference

通过检索增强的可靠性感知推理缓解多模态系统中的视觉幻觉

Pratheswaran Hariharan, Haiping Xu, Donghui Yan

机构 * University of Massachusetts, Dartmouth(马萨诸塞大学达特茅斯分校)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出一种检索增强的可靠性感知推理框架,利用外部视觉证据库和多个可靠性指标进行决策门控,在不重训练模型的情况下减少视觉幻觉,将接受预测准确率从85.84%提升至88.88%。

Comments 29 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07649 2026-07-22 cs.CV cs.AI 版本更新 57%

ViMax: Agentic Video Generation

ViMax: 智能体视频生成

Lingxuan Huang, Sizhe He, Hengji Zhou, Liqiang Nie, Lianghao Xia, Chao Huang

机构 * The University of Hong Kong(香港大学) South China University of Technology(华南理工大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出ViMax框架,通过多智能体协作实现长视频生成,利用分层叙事引擎和视觉一致性机制,保证叙事连贯性和视觉一致性。

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30296 2026-07-02 cs.AI 版本更新 57%

ManimAgent: Self-Evolving Multimodal Agents for Visual Education

ManimAgent: 用于视觉教育的自进化多模态智能体

Wenjia Jiang, Zongyuan Cai, Yuanhang Shao, Chenru Wang, Boyan Han, Zhixue Song, Keyu Chen, Shengwei An, Xu Yang, Zhou Yang

机构 * University of Alberta(阿尔伯塔大学) Southeast University(东南大学) Virginia Tech(弗吉尼亚理工学院) Xidian University(西安电子科技大学) Vivavia Inc(Vivavia公司)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出ManimAgent,通过双通道情节记忆库跨任务传递反思经验,无需权重更新或人工种子,在代码生成任务中提升通过率并减少反思轮次。

Comments Project page: https://manimagent.github.io/. Code: https://github.com/jwj1342/Paper2Manim

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12735 2026-08-06 cs.CV 版本更新 50%

AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition

AffectAgent:协同多智能体推理用于检索增强的多模态情感识别

Zeheng Wang, Zitong Yu, Yijie Zhu, Bo Zhao, Haochen Liang, Taorui Wang, Wei Xia, Jiayu Zhang, Zhishu Liu, Hui Ma, Fei Ma, Qi Tian

机构 * Great Bay University(大湾大学) Guangming Laboratory(光明实验室)

专题命中 多模态RAG :retrieval-augmented generation(abstract)

AI总结 本文提出AffectAgent,通过协同多智能体推理框架,解决多模态情感识别中的模态模糊和复杂情感依赖问题,引入MB-MoE和RAAF提升表现。

Comments Accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02970 2026-08-04 cs.HC 版本更新 50%

From Explanation to Diagnosis: Next Generation Interactive Video Coach with Misstep Awareness

从解释到诊断:具有失误感知能力的下一代交互式视频教练

Xiao Jin, Rahul K. Dass, Ashok K. Goel

专题命中 多模态RAG :knowledge retrieval(abstract)

AI总结 提出一种基于双模型架构的失误感知教练能力,通过将任务-方法-知识模型与教学模型结合,实现学习者错误的检测、分类和诊断性支架生成,从而提供更精确、可操作的反馈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05256 2026-07-03 cs.CV 版本更新 50%

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum

Wiki-R1: 通过数据和采样课程激励基于知识的多模态推理用于VQA

Shan Ning, Longtian Qiu, Xuming He

机构 * ShanghaiTech University(上海科技大学) Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程研究中心) Lingang Laboratory(临港实验室)

专题命中 多模态RAG :retriever(abstract)

AI总结 提出Wiki-R1框架,通过可控课程数据生成和课程采样策略,结合强化学习激励MLLMs在KB-VQA中的推理能力,在Encyclopedic VQA和InfoSeek上取得新SOTA。

Comments Accepted by ICLR 26, code and weights are publicly available

详情

展开后加载摘要…

URL PDF HTML 收藏