arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 546 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 546 篇

2605.27378 2026-05-28 cs.CL cs.CV cs.MA 81%

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

OralAgent: 融合推理、工具与知识的交互式牙科影像分析

Jing Hao, Siyuan Dai, Yongxin Zhang, Yuci Liang, Jiamin Wu, Jiahao Bao, Yuxuan Fan, Zanting Ye, Yanpeng Sun, Xinyu Zhang, Ming Hu, Liang Zhan, James Kit Hon Tsoi, Linlin Shen, Junjun He, Kuo Feng Hung

机构 * Faculty of Dentistry, the University of Hongkong, Hong Kong SAR, China(香港大学牙科学院,中国香港特别行政区) Department of Electrical and Computer Engineering, University of Pittsburgh, Pittsburgh, PA, USA(匹兹堡大学电气与计算机工程系,美国宾夕法尼亚州匹兹堡) Shenzhen University, China(深圳大学,中国) Department of Craniomaxillofacial Surgery, Shanghai Ninth People’s Hospital, China(上海第九人民医院口腔颌面外科部,中国) Nanyang technological University, Singapore(南洋理工大学,新加坡) School of Biomedical Engineering, Southern Medical University, China(南方医科大学生物医学工程学院,中国) Singapore University of Technology and Design, Singapore(新加坡科技设计大学,新加坡) University of Auckland, new zealand(奥克兰大学,新西兰) Shanghai Artificial Intelligence Laboratory , China(上海人工智能实验室,中国)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);knowledge retrieval(abstract);分类 cs.CL

AI总结 提出首个牙科专用AI智能体OralAgent,通过集成22种视觉分析工具和368本经典牙科教科书,实现多模态推理、工具决策与知识检索的自动化框架,在多个基准上达到最优性能。

Comments 14 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21447 2026-02-26 cs.CR cs.AI cs.CL cs.LG 81%

Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG

对抗意图是潜在变量:用于安全多模态代理RAG的状态信任推断

Inderjeet Singh, Vikas Pahuja, Aishvariya Priya Rathina Sabapathy, Chiara Picardi, Amit Giloni, Roman Vainshtein, Andrés Murillo, Hisashi Kojima, Motoyoshi Sekiya, Yuki Unno, Junichi Suga

机构 * Fujitsu Research of Europe, UK(富士通欧洲研究机构,英国) Fujitsu Limited, Japan(富士通株式会社,日本)

专题命中 多模态RAG :RAG(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出MMA-RAG^T框架,通过状态化信任推断提升多模态代理RAG的安全性,实验显示攻击成功率显著降低。

Comments 13 pages, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04790 2025-03-10 cs.CL cs.AI 81%

SuperRAG: Beyond RAG with Layout-Aware Graph Modeling

Jeff Yang, Duy-Khanh Vu, Minh-Tien Nguyen, Xuan-Quang Nguyen, Linh Nguyen, Hung Le

专题命中 多模态RAG :RAG(title,abstract);分类 cs.CL、cs.AI

Comments NAACL 2025, Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23132 2026-07-28 cs.CV 新提交 80%

DispatchRAG: Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video

DispatchRAG:基于交通事故视频中的现实世界协议进行应急调度决策

Muhammad Sulthan Adhipradhana, Ehsan Javanmardi, Naren Bao, Manabu Tsukada

机构 * Graduate School of Information Science and Technology, University of Tokyo(东京大学信息科学与技术研究生院)

专题命中 多模态RAG :RAG(summary_cn,abstract)

AI总结 研究旨在通过DispatchRAG框架,利用基于RAG的检索机制和大型语言模型驱动的推理器,根据日本现实交通事故响应协议进行事故评估和调度,引入事故调度数据集验证框架,为自动驾驶车辆事故报告提供支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31219 2026-06-11 cs.CV cs.CR cs.LG 版本更新 80%

Latent Geometric Chords for Query-Efficient Decision-Based Adversarial Attacks

潜在几何和弦:面向查询高效决策型对抗攻击

Ei Hmue Khine, Yao Li, Jiebao Sun, Shengzhu Shi, Zhichang Guo, Boying Wu

专题命中 多模态RAG :RAG(summary_cn,abstract)

AI总结 提出潜在几何和弦(LGC)方法,通过曲率感知的几何搜索在压缩语义流形中导航决策边界,并引入残差对抗生成(RAG)机制以高视觉保真度实现查询高效的决策型黑盒对抗攻击。

Comments Added a conceptual diagram for the LGC architecture, 14 pages, 10 figures, 7 tables. Submitted to IEEE Transactions on Information Forensics and Security. The source code is available at https://github.com/eihmuekhine/Latent-Geometric-Chords

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16313 2026-04-21 cs.IR cs.AI cs.CL 80%

MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering

MARA:一种多模态自适应检索增强框架用于文档问答

Hui Wu, Haoquan Zhai, Yuchen Li, Hengyi Cai, Peirong Zhang, Yidan Zhang, Lei Wang, Chunle Wang, Yingyan Hou, Shuaiqiang Wang, Dawei Yin

机构 * Key Laboratory of Target Cognition and Application Technology (TCAT), AIRI, CAS(目标认知与应用技术重点实验室(TCAT),空气动力研究所,中国科学院) School of Electronic, Electrical and Communication Engineering, UCAS(电子、电气与通信工程学院,中国科学院大学) Aerospace Information Research Institute, Chinese Academy of Sciences(航天信息研究所,中国科学院) Baidu Inc.(百度公司)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 MARA框架通过引入查询自适应机制提升多模态文档检索与生成的精度和质量,实验表明其在多模态问答基准上优于现有最先进方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09733 2026-07-29 cs.CL cs.CV 版本更新 79%

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation

VisRAG2.0:通过视觉检索增强生成中的证据引导多图像推理减轻视觉幻觉

Yubo Sun, Chunyi Peng, Yukun Yan, Shi Yu, Zhenghao Liu, Sen Mei, Chi Chen, Maosong Sun

机构 * School of Software and Microelectronics, Peking University, China(北京大学软件与微电子学院) School of Computer Science and Engineering, Northeastern University, China(东北大学计算机科学与工程学院) Department of Computer Science and Technology, Institute for AI, Tsinghua University, China(清华大学人工智能研究院计算机科学与技术系)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);分类 cs.CL

AI总结 研究针对视觉检索增强生成中VLM存在的视觉幻觉及证据识别问题,提出证据引导的多图像推理框架EVisRAG,引入RS-GRPO改进训练,实验表明该方法能提升性能、减少幻觉,有效提高视觉基础和推理可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22643 2026-07-28 cs.AI cs.CV 新提交 79%

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

检索前先思考:多模态检索增强生成的智能体规划

Tianyu Yang, Shir Simon, Zhenzhen Li, Minhao Cheng, Xiangliang Zhang

机构 * Bosch AI Research Center(博世人工智能研究中心) University of Notre Dame(圣母大学) Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 多模态RAG :RAG(title);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 研究多模态检索增强生成问题,提出MM-R2框架,通过建模检索内容和位置在检索前推理,构建意图基础检索状态,在结构化知识图谱检索,还构建相关数据集并采用两阶段后训练策略,实验证明其在答案准确性等方面表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17644 2026-01-28 cs.CR cs.AI 79%

A Systemic Evaluation of Multimodal RAG Privacy

多模态RAG隐私的系统评估

Ali Al-Lawati, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 多模态RAG :RAG(title);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文系统评估了多模态RAG隐私风险,揭示了推理过程中可能泄露私有信息的问题,并呼吁开发隐私保护机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10337 2026-01-15 cs.AI cs.LG 79%

A Curriculum Learning Approach to Reinforcement Learning: Leveraging RAG for Multimodal Question Answering

一种基于RAG的强化学习课程学习方法:用于多模态问答

Chenliang Zhang, Lin Wang, Yuanyuan Lu, Yusheng Qi, Kexin Wang, Peixu Hou, Wenshi Chen

机构 * Meituan(美团)

专题命中 多模态RAG :RAG(title);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出了一种结合课程学习与强化学习的方法,用于多模态问答任务,通过检索增强生成系统在挑战中取得优异成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08226 2026-01-14 cs.CV cs.AI 79%

Knowledge-based learning in Text-RAG and Image-RAG

基于知识的学习在Text-RAG和Image-RAG中的应用

Alexander Shim, Khalil Saieh, Samuel Clarke

机构 * Florida International University(佛罗里达国际大学)

专题命中 多模态RAG :RAG(title,abstract);分类 cs.AI

AI总结 本研究通过比较基于文本和图像的RAG方法,探讨了如何利用外部知识减少幻觉问题并提升胸部X光图像疾病检测的准确性。

Comments 9 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18987 2025-12-23 cs.RO cs.CL cs.CV 79%

Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation

语义可感知的多模态检索:基于具身记忆的层次化移动操作

Ryosuke Korekata, Quanting Xie, Yonatan Bisk, Komei Sugiura

机构 * Keio University(keio大学) Keio AI Research Center(keio人工智能研究中心) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态RAG :RAG(title,abstract);分类 cs.CL

AI总结 本研究提出Affordance RAG框架,通过构建具有可操作性的具身记忆,提升机器人在开放词汇移动操作中的检索性能和任务成功率。

Comments Accepted to IEEE RA-L, with presentation at ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21002 2025-11-27 cs.CV cs.AI 79%

Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning

知识完善视觉:一种多模态实体感知检索增强生成框架用于新闻图像描述

Xiaoxing You, Qiang Huang, Lingyu Li, Chi Zhang, Xiaopeng Liu, Min Zhang, Jun Yu

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);分类 cs.AI

AI总结 MERGE提出了一种多模态实体感知检索增强生成框架,通过构建实体中心知识库和改进跨模态对齐,提升新闻图像描述质量和命名实体识别性能。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09266 2025-10-13 cs.CL 79%

CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation

Kaiwen Wei, Xiao Liu, Jie Zhang, Zijian Wang, Ruida Liu, Yuming Yang, Xin Xiao, Xiao Sun, Haoyang Zeng, Changzai Pan, Yidan Zhang, Jiang Zhong, Peijin Wang, Yingchao Feng

机构 * Chongqing University(重庆大学) Independent Researcher(独立研究者) University of the Chinese Academy of Sciences(中国科学院大学) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11937 2025-09-16 cs.SE cs.AI 79%

MMORE: Massive Multimodal Open RAG & Extraction

Alexandre Sallinen, Stefan Krsteski, Paul Teiletche, Marc-Antoine Allard, Baptiste Lecoeur, Michael Zhang, Fabrice Nemo, David Kalajdzic, Matthias Meyer, Mary-Anne Hartley

机构 * École Polytechnique Fédérale de Lausanne (EPFL), Switzerland(瑞士联邦理工学院洛桑校区) ETH Zürich, Switzerland(瑞士苏黎世联邦理工学院) T.H. Chan School of Public Health, Harvard University, USA(哈佛大学T.H. Chan公共卫生学院)

专题命中 多模态RAG :RAG(title,abstract);分类 cs.AI

Comments This paper was originally submitted to the CODEML workshop for ICML 2025. 9 pages (including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14864 2025-02-21 cs.AI cs.CV 79%

Benchmarking Multimodal RAG through a Chart-based Document Question-Answering Generation Framework

Yuming Yang, Jiang Zhong, Li Jin, Jingwang Huang, Jingpeng Gao, Qing Liu, Yang Bai, Jingyuan Zhang, Rui Jiang, Kaiwen Wei

专题命中 多模态RAG :RAG(title);retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10834 2025-01-22 cs.CV cs.AI cs.LG 79%

Visual RAG: Expanding MLLM visual knowledge without fine-tuning

Mirco Bonomo, Simone Bianco

专题命中 多模态RAG :RAG(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06962 2025-06-17 cs.CV 79%

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation

Jingyuan Qi, Zhiyang Xu, Qifan Wang, Lifu Huang

机构 * Virginia Tech(弗吉尼亚理工大学) Meta UC Davis(加州大学戴维斯分校)

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(comments)

Comments Image Generation, Retrieval Augmented Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13706 2026-08-17 cs.CL cs.AI 新提交 79%

CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA

CLAIR-Fin:用于跨模态金融问答中声明级验证与自适应辩论的对抗性多智能体框架

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain

机构 * Ahsanullah University of Science and Technology(阿萨努拉科技大学) Jashore University of Science and Technology(杰索尔科技大学) American International University - Bangladesh(孟加拉国美国国际大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出CLAIR-Fin九智能体框架,针对跨模态金融问答的声明级验证与自适应辩论,在BB-FinQA-X数据集上提升了模型忠实度,且弃权比例合理,优于相关基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03836 2026-07-07 cs.CV cs.AI cs.CL 新提交 79%

When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

何时越简单越好:评估中世纪拉丁文手稿的翻译管道

Nguyen Kim Hai Bui, Md. Easin Arafat, Tamás Gábor Orosz, Mufti Mahmud

机构 * Eötvös Loránd University(厄特沃什·罗兰大学) King Fahd University of Petroleum and Minerals(法赫德国王石油与矿产大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 针对历史手稿翻译难题,提出评估中世纪拉丁文手稿图像到翻译流程的框架。通过CATMuS数据集对比发现领域特定OCR模型优势,引入新数据集IPC,实验揭示简单管道表现更佳,为低资源历史场景翻译系统部署提供基准与指导。

Comments 17 pages, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28607 2026-05-28 cs.AI cs.CL 79%

Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution

基于自适应多智能体框架的自动工作流执行

Susanna Cifani, Mario Luca Bernardi, Marta Cimitile

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) Department of Engineering University of Sannio(萨尼奥大学工程系) Faculty of Jurisprudence Unitelma Sapienza University(法理学院萨皮恩扎大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 提出一种多模态多智能体框架,通过离线构建拓扑知识库和在线自适应检索增强生成与闭环协作验证,实现自动工作流执行。

Comments Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses. Accepted for publication at the 2026 IEEE International Conference on Evolving and Adaptive Intelligent Systems (EAIS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18774 2026-05-20 cs.IR cs.AI 79%

M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language Models

M3DocDep: 多模态、多页、多文档依赖分块方法基于大视觉-语言模型

Joongmin Shin, Jeongbae Park, Jaehyung Seo, Heuiseok Lim

机构 * Human-inspired AI Research, Korea University(韩国大学人智AI研究所) Computer Science and Engineering, Konkuk University(konkuk大学计算机科学与工程系) Department of Computer Science and Engineering, Korea University(韩国大学计算机科学与工程系)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文提出M3DocDep,一种基于大视觉-语言模型的多模态、多页、多文档依赖分块方法,通过恢复块级依赖并构建分块,提高了长多页多模态文档的检索和问答质量。

Comments Accepted to CVPR2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16275 2026-05-19 cs.CY cs.AI cs.CL cs.MM 79%

AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course

AI 产出物还是AI增强?英语学术用途课程中学生对AI生成媒体的看法

David James Woo, Deliang Wang, Kai Guo

机构 * Everwrite Limited(Everwrite有限公司) Faculty of Education, The University of Hong Kong(香港大学教育学院) Faculty of Education, The Chinese University of Hong Kong(香港中文大学教育学院)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了AI生成内容在EAP课程中的教学效果,通过混合方法分析发现学生偏好视觉化内容,视频与学业表现正相关,但高认知负荷与成绩负相关,表明需合理设计内容以提升学习效果。

Comments 23 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11864 2026-05-13 cs.IR cs.AI cs.CV cs.MM 79%

Very Efficient Listwise Multimodal Reranking for Long Documents

非常高效的长文档多模态重排序方法

Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh

机构 * Magellan Technology Research Institute (MTRI)(马杰拉技术研究院(MTRI))

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文提出ZipRerank,通过轻量级查询-图像早期交互机制和单次前向传递消除自回归解码,实现高效多模态重排序,实验表明其在MMDocIR基准上性能优异且显著降低LLM推理延迟。

Comments To appear in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07241 2026-07-10 cs.SD 新提交 78%

Rag Classification of Tagore Songs using Symbolic Music Notation and Novel Weighted Distance Measures

使用符号音乐记谱法和新型加权距离度量对泰戈尔歌曲进行拉格分类

Chandan Misra, Swarup Chattopadhyay

机构 * XIM University(西姆大学)

专题命中 多模态RAG :RAG(title,abstract)

AI总结 该研究针对罗宾德拉·桑吉特歌曲拉格识别难题,将其转化为监督分类问题,利用符号乐谱记谱法构建数据集,探讨多种距离度量,引入加权欧几里得距离,在k近邻框架下改进拉格分类,更好捕捉旋律特征。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20414 2026-05-26 eess.AS 78%

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding

PlanRAG-Audio:面向长音频理解的规划与检索增强生成

Masao Someki, Chien-yu Huang, Siddhant Arora, Samuele Cornell, Markus Müller, Nathan Susanj, Rupak V Swaminathan, Grant P Strimel, Jing Liu, Shinji Watanabe

专题命中 多模态RAG :retrieval augmented generation(title);retrieval-augmented generation(abstract)

AI总结 提出PlanRAG-Audio框架,通过规划查询所需模态和时间跨度并仅检索相关信息,实现长音频的高效推理,提升准确率并稳定性能。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04372 2026-04-07 cs.CV 78%

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning

图到帧RAG:面向无训练和可审计的视频推理的视觉空间知识融合

Songyuan Yang, Weijiang Yu, Ziyu Liu, Guijian Tang, Wenjing Yang, Huibin Tan, Nong Xiao

机构 * National University of Defense Technology(国防科技大学) Sun Yat-sen University(中山大学)

专题命中 多模态RAG :RAG(title,abstract)

AI总结 本文提出G2F-RAG,通过视觉空间知识融合提升视频推理的可解释性和效率,减少认知负担并保留可追溯的证据轨迹。

Comments Accepted at CVPR 2026. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02258 2026-03-24 cs.CV 78%

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

病理代理RAG:通过强化学习实现多模态代理检索增强生成用于病理学视觉语言模型

Wenchuan Zhang, Jingru Guo, Hengzhe Zhang, Penghao Zhang, Jie Chen, Shuwan Zhang, Zhang Zhang, Yuhao Yi, Hong Bu

专题命中 多模态RAG :retrieval-augmented generation(title);RAG(abstract)

AI总结 本文提出Patho-AgenticRAG,通过强化学习实现多模态代理检索增强生成,解决病理学视觉语言模型在高分辨率、复杂组织结构和临床语义上的挑战,提升诊断准确性。

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(35): 29921-29929, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23483 2025-12-30 cs.CV 78%

TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding

TV-RAG:一种具有时间意识和语义熵权的长视频检索与理解框架

Zongsheng Cao, Yangfan He, Anran Liu, Feng Chen, Zepeng Wang, Jun Xie

机构 * Researcher(研究者)

专题命中 多模态RAG :RAG(title,abstract)

AI总结 TV-RAG通过时间衰减检索和熵加权关键帧采样,提升长视频检索与理解性能,无需重新训练即可集成至现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05872 2025-12-10 cs.CV 78%

Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection

领域-RAG:跨领域少样本目标检测的检索引导组合图像生成

Yu Li, Xingyu Qiu, Yuqian Fu, Jie Chen, Tianwen Qian, Xu Zheng, Danda Pani Paudel, Yanwei Fu, Xuanjing Huang, Luc Van Gool, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Fuzhou University(福州大学) East China Normal University(华东师范大学) HKUST(GZ)(香港科技大学(广州))

专题命中 多模态RAG :RAG(title,abstract)

AI总结 Domain-RAG通过检索引导的组合图像生成方法,解决跨领域少样本目标检测中的领域对齐与背景生成问题,实现高质量样本生成。

详情

展开后加载摘要…

URL PDF HTML 收藏