arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 544 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 544 篇

2607.24748 2026-07-29 cs.IR cs.AI cs.CL 新提交 93%

VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents

VLD-RAG:用于长的、视觉丰富的多页文档的智能视觉语言检索增强生成

Seonok Kim

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(title,abstract);hybrid retrieval(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 研究针对视觉丰富的长文档问答的多模态检索增强生成,提出VLD-RAG框架,构建多模态索引并采用混合检索策略,经智能体工作流程协调,在相关基准上提升了证据页面检索及问答表现,凸显协调验证与混合检索的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04969 2026-07-14 cs.IR cs.AI 版本更新 93%

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation

MG$^2$-RAG:多粒度图用于多模态检索增强生成

Sijun Dai, Qiang Huang, Xiaoxing You, Jun Yu

机构 * School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机学院)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(title,abstract);分类 cs.IR、cs.AI

AI总结 MG$^2$-RAG通过构建层次化多模态知识图谱,融合文本与视觉信息,提升多模态检索与生成性能,实验显示其在四个任务中均取得最佳效果,同时显著提升效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06185 2026-05-08 cs.AI cs.CV 92%

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios

事件因果RAG:一种用于长视频复杂场景中长视频推理的检索增强生成框架

Peizheng Yan, Yu Zhao, Liang Xie, Juntong Qi, Mingming Wang, Erwei Yin

机构 * Tianjin Key Lab of Intelligent Unmanned Swarm Tech & System(天津智能集群技术与系统重点实验室) Tianjin University(天津大学) Institute of Computing and Intelligence(计算与智能研究所) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Tianjin Artificial Intelligence Innovation Center(天津人工智能创新中心) Defense Innovation Institute Academy of Military Sciences(国防科技创新研究院) School of Future Technology(未来技术学院) Shanghai University(上海大学)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(title,abstract);分类 cs.AI

AI总结 本文提出Event-Causal RAG框架,通过事件因果图和双存储记忆实现长视频推理,优于传统方法,在多事件整合和因果推理任务中表现更优,同时提升内存效率和流式性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09331 2026-08-11 cs.SD cs.AI 新提交 92%

RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction

RAG-Audio:用于忠实脑电信号到音频重建的检索增强生成

Ambuj Mehrish, Sebastiano Vascon

机构 * CVML Lab(CVML实验室)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(title);分类 cs.AI

AI总结 RAG-Audio通过检索真实音频样本初始化生成器采样轨迹,缓解脑电到音频生成的先验主导问题,在Brain2Music数据集上提升了刺激识别准确率并大幅降低Fréchet音频距离。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28570 2026-06-30 cs.CV cs.AI cs.MA 92%

Digitizing Coaching Intelligence: An Agentic Framework for Holistic Athlete Profiling using VLM and RAG

数字化教练智能:基于VLM和RAG的整体运动员画像智能体框架

Deep Ghosal, Ishani Sen, Wazib Ansar, Amlan Chakrabarti

机构 * A.K. Choudhury School of Information Technology University of Calcutta(A.K.楚德里信息科技学院印度加尔各答大学)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);vector search(abstract);分类 cs.AI

AI总结 提出基于LLM的混合智能体框架,结合计算机视觉与视觉语言模型,通过3×3智能网格分块策略降低计算开销,并引入LLM-as-a-Judge自校正循环和双持久化RAG管道,实现符合印度体育局标准的自动化整体运动员评估。

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19986 2026-06-24 cs.IR cs.CV 92%

Automating Iconclass: LLMs and RAG for Large-Scale Classification of Religious Woodcuts

自动化Iconclass:LLM和RAG用于大规模宗教木刻版画分类

Drew B. Thomas

机构 * University College Dublin(都柏林大学)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);vector search(abstract);分类 cs.IR

AI总结 本文提出利用LLM和向量数据库结合RAG技术,实现早期现代宗教图像的高效分类,通过全页上下文生成描述并匹配Iconclass代码,提升分类精度至92%。

Comments 29 pages, 7 figures. First presented at the "Digital Humanities and Artificial Intelligence" conference at the University of Reading on 17 June 2024

Journal ref Digital Culture and Education 16(3):161-182, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12121 2026-08-13 cs.CL cs.AI 新提交 91%

QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving

QV-PIC:面向高效RAG服务的查询感知视觉位置无关缓存

Yilin Liu, Rui Meng, Wangze Ni, Jianxin Yan, Heng Cao, Libin Zheng, Peng Cheng, Jinfei Liu

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 QV-PIC是一种查询感知双分辨率PIC复用框架,通过离线编译视觉缓存、在线分分辨率处理,解决渲染图像PIC质量下降问题,在6项RAG任务中显著提升F1并降低TTFT。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14445 2026-06-23 cs.CL cs.AI cs.HC cs.LG 版本更新 91%

Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning

Tell Me:基于LLM的心理健康助手,集成RAG、合成对话生成与智能体规划

Trishala Jayesh Ahalpara

机构 * Fujitsu Research of America(富士通美国研究)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 提出Tell Me系统,利用大语言模型通过RAG实现个性化对话、合成客户-治疗师对话生成以及智能体规划生成自我护理计划,旨在提供可及的心理支持而非替代专业治疗。

Comments 8 pages, 2 figures, 1 Table. Submitted to the Computation and Language (cs.CL) category. Uses the ACL-style template. Code and demo will be released at: https://github.com/trystine/Tell_Me_Mental_Wellbeing_System

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28580 2026-07-31 cs.AI 新提交 91%

DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation

DualG-MRAG:面向多模态检索增强生成的宏观推理与微观匹配解耦框架

Jiacheng Tao, Qingyun Sun, Haonan Yuan, Ziwei Zhang, Jianxin Li

机构 * Beihang University(北京航空航天大学) SKLCCSE

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(title,abstract);retriever(abstract,abstract_cn);分类 cs.AI

AI总结 针对多模态RAG在复杂多跳推理中存在的问题,本文提出DualG-MRAG框架,通过解耦宏观推理与微观匹配抑制检索噪声,经实验验证其在证据召回和问答准确率上优于基线。

Comments Accepted to the 34th ACM International Conference on Multimedia (ACM MM 2026). 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22378 2026-07-14 cs.SD cs.AI cs.MM eess.AS 91%

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach

零努力图像到音乐生成:一种可解释的基于RAG的视觉语言模型方法

Zijian Zhao, Dian Jin, Zijing Zhou

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) The Hong Kong Polytechnic University(香港理工大学) The University of Hong Kong(香港大学) The Hong Kong University of Science(香港科学大学) The Hong Kong Polytechnic University Hong Kong China(香港理工大学香港中国) The University of Hong Kong Hong Kong China(香港大学香港中国)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出一种基于RAG的视觉语言模型方法,实现图像到音乐生成,通过ABC记谱法连接文本与音乐模态,并利用多模态检索增强生成和自反思技术,提供高可解释性且低计算成本的音乐生成方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26793 2026-06-26 cs.CR cs.AI cs.LG 新提交 91%

MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

MIRROR: 基于新颖性约束的记忆引导MCTS红队测试用于智能体RAG

Inderjeet Singh, Andrés Murillo, Motoyoshi Sekiya, Yuki Unno, Junichi Suga

机构 * Fujitsu Research of Europe, United Kingdom(欧洲富士通研究机构,英国) Fujitsu Limited, Japan(日本富士通有限公司)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出MIRROR框架,通过记忆引导蒙特卡洛树搜索和显式新颖性约束,在多模态智能体RAG系统的四个攻击面上实现高攻击成功率,并降低跨表面方差。

Comments 6 pages, 2 figures. Accepted at the 2026 International Joint Conference on Neural Networks (IJCNN 2026), IEEE WCCI 2026; presented as an oral talk. Code and ART-SafeBench benchmark: https://github.com/FujitsuResearch/mirror

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15906 2026-06-16 cs.IR cs.AI cs.CL cs.DB cs.MM 新提交 91%

MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA

MAGE-RAG:面向长文档问答的多粒度自适应图证据多模态RAG

Yilong Zuo, Xunkai Li, Jing Yuan, Qiangqiang Dai, Hongchao Qin, Ronghua Li

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态RAG :RAG(title,title_cn);分类 cs.IR、cs.CL、cs.AI

AI总结 提出MAGE-RAG框架,通过离线构建包含页面和元素节点的证据图,在线自适应构建证据子图,平衡证据覆盖与噪声控制,在长文档多模态问答中取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17832 2026-05-28 cs.LG cs.AI cs.CR cs.CV 91%

MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks

MM-PoisonRAG:通过局部和全局投毒攻击破坏多模态RAG

Hyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim, Dimitrios Bralios, Saikrishna Sanniboina, Nanyun Peng, Kai-Wei Chang, Daniel Kang, Heng Ji

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California Los Angeles(加州大学洛杉矶分校)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出MM-PoisonRAG框架,通过局部投毒攻击(LPA)和全局投毒攻击(GPA)两种策略,系统研究多模态检索增强生成(RAG)在知识投毒下的脆弱性,实验表明攻击成功率高达56%且能绕过现有防御。

Comments Code is available at https://github.com/HyeonjeongHa/MM-PoisonRAG

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15019 2026-05-15 cs.CL 91%

From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG

从场景到元素:面向可验证多模态RAG的多粒度证据检索

Guanhua Chen, Chuyue Huang, Yutong Yao, Shudong Liu, Xueqing Song, Lidia S. Chao, Derek F. Wong

机构 * NLP 2 CT Lab, Department of Computer and Information Science, University of Macau(NLP2CT实验室,计算机与信息科学系,澳门大学)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.CL

AI总结 本文提出GranuRAG框架,通过多粒度跨模态对齐和元素级检索,解决多模态RAG中粗粒度证据与细粒度查询不匹配的问题,提升可验证性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08133 2026-05-13 cs.CV cs.AI 91%

VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving

VLADriver-RAG:基于检索增强的视觉-语言-动作模型用于自动驾驶

Rui Zhao, Haofeng Hu, Zhenhai Gao, Jiaqiao Liu, Gao Fei

机构 * College of Automotive Engineering(汽车工程学院) The National Key Laboratory of Automotive Chassis Integration and Bionics(汽车底盘集成与生物力学国家级重点实验室) ReeFocus AI Technology(ReeFocus人工智能技术)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出VLADriver-RAG,通过构建结构化的历史知识,提升自动驾驶中长尾场景的泛化能力,采用视觉到场景机制和场景对齐嵌入模型,结合查询驱动的VLA骨干网络,实现高精度轨迹合成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06173 2026-05-11 cs.CV cs.AI 91%

Retina-RAG: Retrieval-Augmented Vision-Language Modeling for Joint Retinal Diagnosis and Clinical Report Generation

Retina-RAG:基于检索增强的视觉-语言模型用于联合视网膜诊断和临床报告生成

Abdelrahman Zaian, Sheethal Bhat, Mohamed Abdalkader, Andreas Maier

机构 * Friedrich-Alexander-Universität(弗里德里希-亚历山大大学)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 Retina-RAG通过整合高精度视网膜分类器和参数高效视觉语言模型,实现糖尿病视网膜病变分级、黄斑水肿检测及报告生成,优于现有方法,且在单块消费级GPU上运行。

Comments 10 pages, 5 figures. Submitted to MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07273 2026-05-11 cs.CV cs.AI 91%

From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG

从云层到幻觉:遥感视觉-语言RAG中的大气检索劫持

Jiaju Han, Chao Li, Chengyin Hu, Qike Zhang, Xuemeng Sun, Xin Wang, Fengyu Zhang, Xiang Chen, Yiwei Wei, Jiahuan Long, Jiujiang Guo

机构 * China University of Petroleum, Beijing at Karamay(中国石油大学(北京)克拉玛依校区)

专题命中 多模态RAG :RAG(title,title_cn);retriever(abstract);分类 cs.AI

AI总结 研究提出CloudWeb攻击,通过修改输入图像实现遥感多模态RAG中的大气证据劫持,展示了大气变化对检索阶段的影响,并揭示了自然外观的天气变化可能破坏证据检索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27600 2026-05-01 cs.IR 91%

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG

净化多模态检索:用于RAG的片段级证据选择

Xihang Wang, Zihan Wang, Chengkai Huang, Cao Liu, Ke Zeng, Quan Z. Sheng, Lina Yao

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.IR

AI总结 本文提出FES-RAG框架,通过片段级证据选择提升多模态检索效果,减少噪声干扰,实验显示在M2RAG基准上性能提升27%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14967 2026-04-20 cs.CV cs.AI 91%

UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

UniDoc-RL: 基于分层动作和密集奖励的粗到细视觉RAG

Jun Wang, Shuo Tan, Zelong Sun, Tiancheng Gu, Yongle Zhao, Ziyong Feng, Kaicheng Yang, Zhiwu Lu

机构 * DeepGlint-AI

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 UniDoc-RL通过分层动作空间和密集奖励方案,提升视觉RAG系统的细粒度视觉语义处理能力,实现端到端训练,实验显示优于现有方法。

Comments 17 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21956 2025-09-30 cs.CV cs.AI cs.CL cs.LG 91%

Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation

Mengdan Zhu, Senhao Cheng, Guangji Bai, Yifei Zhang, Liang Zhao

机构 * Emory University(埃默里大学) University of Michigan, Ann Arbor(密歇根大学安娜堡分校)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(title,abstract);retriever(abstract);hybrid retrieval(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16604 2026-07-21 eess.IV 新提交 90%

When Do Multimodal and Graph-Augmented RAG Help? A Controlled Evaluation for Document Question Answering

多模态和图增强的检索增强生成(RAG)何时有用?文档问答的对照评估

Sokipriala Jonah

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract)

AI总结 研究文档问答中多模态和图增强RAG的作用,提出用多模态图-RAG架构,通过独立检索并融合文本、图和视觉证据源来评估效果。经实验发现其价值取决于多种因素,如检索设计等,还揭示了一些相关问题,如图像检索及生成器效率对结果的影响。

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05868 2026-07-08 cs.CR 新提交 90%

Code-Level Cost Function Generation for Spatial Image Steganography Using RAG-Enhanced Large Language Models

使用RAG增强大语言模型进行空间图像隐写术的代码级成本函数生成

Yige Wang, Shiqi Yi, Hanzhou Wu

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract)

AI总结 该研究针对空间图像隐写术成本函数设计难题,提出利用RAG增强大语言模型的进化系统,含SE-RAG模块及反馈机制。实验表明,此框架安全性更高,还提升了代码执行率并降低搜索成本,凸显结合大语言模型与领域知识的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02553 2026-06-02 cs.CV 90%

LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

LongLive-RAG: 一种用于长视频生成的通用检索增强框架

Qixin Hu, Shuai Yang, Wei Huang, Song Han, Yukang Chen

机构 * NVIDIA USC(美国大学) MIT(麻省理工学院)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract)

AI总结 提出LongLive-RAG框架,通过将自回归视频生成中的历史潜变量作为可检索记忆,利用查询嵌入检索相关历史潜变量并引入窗口时间增量损失,以减轻滑动窗口注意力导致的误差累积,提升长视频生成质量。

Comments 20 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18884 2026-05-20 cs.LG cs.CV 90%

Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition

在情绪树中导航:用于多模态情绪识别的分层双曲RAG

Zeheng Wang, Bo Zhao, Yijie Zhu, Zhishu Liu, Hui Ma, Ruixin Zhang, Shouhong Ding, Qianyu Xie, Zitong Yu

机构 * Great Bay University(广东东莞大亚湾大学) Tencent Youtu Lab(腾讯优图实验室)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract)

AI总结 本文提出HyperEmo-RAG,一种利用结构化情绪知识库的检索增强生成框架,通过双曲空间嵌入和证据图构建来提升多模态情绪识别的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12352 2026-04-15 cs.AI cs.CL 90%

MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents

多文档融合:一种用于长工业文档增强RAG的分块流程

Joongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo, Heuiseok Lim

机构 * Human-inspired AI Research(人机协同人工智能研究所) Department of Computer Science and Engineering(计算机科学与工程系) School of Software(软件学院)

专题命中 多模态RAG :RAG(title,title_cn);分类 cs.CL、cs.AI

AI总结 针对长工业文档结构复杂的问题,提出MultiDocFusion分块流程,结合视觉解析、OCR提取、层次结构重建和DFS分组,提升RAG检索精度和问答质量。

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18917 2026-07-22 cs.CV 新提交 89%

TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering

TAP-RAG:用于长文档多模态问答的任务感知策略控制

Zhong Ji, Keqi Jin, Yan Zhang, Jiasheng Li

机构 * School of Electrical and Information Engineering, Tianjin University(天津大学电气与信息工程学院) The International Joint Institute of Tianjin University(天津大学国际联合学院) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态RAG :RAG(title,title_cn)

AI总结 研究长文档多模态问答,提出TAP-RAG框架,含任务感知策略控制器及两个执行器,能预测任务先验、估计证据信号并生成策略,在多模态文档图上扩展证据,在相关数据集上取得最佳总体准确率。

Comments 18 pages, 6 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05438 2026-07-08 cs.IR cs.AI 新提交 89%

Modality Relevance is not Modality Utility: Post-hoc Selective Modality Escalation for Cost-Aware Multimodal RAG

模态相关性并非模态效用:用于成本感知多模态RAG的事后选择性模态升级

Xue Li, Yiming Gai

机构 * Hangzhou International Innovation Institute, Beihang University(杭州国际创新研究院,北京航空航天大学)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 研究多模态检索增强生成中模态选择决策点问题,提出事后选择性模态升级方法,先低成本从文本和表格回答,再按需支付VLM证据费用,校准路由器决定是否升级,在MultiModalQA上减少视觉调用并缩小与神谕升级率差距,扩展了路由信号层次结构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01115 2026-07-02 cs.CL cs.AI 新提交 89%

Towards Developing a Multimodal Chat Assistant for University Stakeholders: RAG-based Approach

面向大学利益相关者的多模态聊天助手开发:基于RAG的方法

Md Abu Hanif Shaikh, Abdullah Al Shafi

机构 * Institute of Information and Communication Technology, Khulna University of Engineering & Technology(信息与通信技术研究所,库尔纳工程技术大学)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 针对大学利益相关者获取信息困难的问题,提出基于检索增强生成的多模态聊天系统,结合大语言模型与语义检索,支持文本和图像查询,将幻觉率从31.7%降至6.6%。

Comments Accepted at 2025 28th International Conference on Computer and Information Technology (ICCIT)

Journal ref 2025 28th International Conference on Computer and Information Technology (ICCIT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01613 2026-06-16 cs.IR cs.AI cs.MA 版本更新 89%

TechRAG: Evidence-Gated Multimodal Agentic RAG for Technical Literature Reasoning

TechGraphRAG:面向技术文献推理的智能图增强RAG框架

Kanwar Bharat Singh

机构 * Global Tire Intelligence and Solutions (GTIS)(全球轮胎智能与解决方案(GTIS)) The Goodyear Tire & Rubber Company(固特异轮胎与橡胶公司)

专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 提出一种13步自主流水线的智能检索增强生成框架,通过证据充分性评分、知识图谱遍历和自校正生成,支持领域特定技术文献推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07924 2026-06-09 cs.CV cs.AI cs.CL cs.LG cs.MM 新提交 89%

Decoupling Semantics and Logic: A Training-Free Coarse-to-Fine Pipeline for Video Retrieval-Augmented Generation

解耦语义与逻辑:一种无需训练的从粗到精的视频检索增强生成流水线

Jiaxin Dai, Zehang Wei, Jiamin Yan, Xiang Xiang

机构 * School of Computer Science & Tech, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) School of AI and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(title);dense retrieval(abstract);分类 cs.CL、cs.AI

AI总结 提出一种无需训练的两阶段级联视频RAG流水线,通过解耦语义检索与逻辑推理,实现跨语言长视频理解、严格角色遵循和零幻觉时间定位。

Comments To be presented at ACL 2026 MAGMAR Workshop (Oral; Retrieval leaderboard No.1)

详情

展开后加载摘要…

URL PDF HTML 收藏