arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 544 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 544 篇

2506.07600 2025-06-10 cs.CV cs.AI 83%

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

Nianbo Zeng, Haowen Hou, Fei Richard Yu, Si Shi, Ying Tiffany He

机构 * Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室) College of Computer Science and Software Engineering(计算机科学与软件工程学院)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07399 2025-06-10 cs.CV cs.AI 83%

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems

Peiru Yang, Jinhua Yin, Haoran Zheng, Xueying Bai, Huili Wang, Yufei Sun, Xintian Li, Shangguang Wang, Yongfeng Huang, Tao Qi

机构 * Tsinghua University(清华大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01222 2025-05-23 cs.CV cs.CL 83%

Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG

Wenbin Wang, Yongcheng Jing, Liang Ding, Yingjie Wang, Li Shen, Yong Luo, Bo Du, Dacheng Tao

机构 * Wuhan University(武汉大学) Nanyang Technological University(南洋理工大学) The University of Sydney(悉尼大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13828 2025-05-21 cs.AI 83%

Multimodal RAG-driven Anomaly Detection and Classification in Laser Powder Bed Fusion using Large Language Models

Kiarash Naghavi Khanghah, Zhiling Chen, Lela Romeo, Qian Yang, Rajiv Malhotra, Farhad Imani, Hongyi Xu

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

Comments ASME 2025 International Design Engineering Technical Conferences and Computers and Information in Engineering Conference IDETC/CIE2025, August 17-20, 2025, Anaheim, CA (IDETC2025-168615)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12663 2025-03-18 cs.CV cs.CL cs.LG cs.RO 83%

Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding

Imran Kabir, Md Alimoor Reza, Syed Billah

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13085 2025-03-04 cs.LG cs.CL cs.CV 83%

MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models

Peng Xia, Kangyu Zhu, Haoran Li, Tianze Wang, Weijia Shi, Sheng Wang, Linjun Zhang, James Zou, Huaxiu Yao

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00239 2024-12-03 cs.SE cs.AI 83%

Generating a Low-code Complete Workflow via Task Decomposition and RAG

Orlando Marquez Ayala, Patrice Béchard

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

Comments Under review; 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11321 2024-10-16 cs.CL 83%

Self-adaptive Multimodal Retrieval-Augmented Generation

Wenjia Zhai

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12309 2024-08-20 cs.CV cs.IR cs.LG 83%

iRAG: Advancing RAG for Videos with an Incremental Approach

Md Adnan Arefeen, Biplob Debnath, Md Yusuf Sarwar Uddin, Srimat Chakradhar

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR

Comments Accepted in CIKM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07016 2024-02-13 cs.AI 83%

REALM: RAG-Driven Enhancement of Multimodal Electronic Health Records Analysis via Large Language Models

Yinghao Zhu, Changyu Ren, Shiyun Xie, Shukai Liu, Hangyuan Ji, Zixiang Wang, Tao Sun, Long He, Zhoujun Li, Xi Zhu, Chengwei Pan

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28002 2026-06-29 cs.CL cs.AI eess.AS 新提交 82%

Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection

对话到检测:用于保险欺诈检测的多模态混合 NLP 流水线

Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth, Karishma Jaitly

机构 * Aston University(阿斯顿大学) Domestic & General

专题命中 多模态RAG :RAG(summary_cn,abstract);分类 cs.CL、cs.AI

AI总结 提出一种合成多模态框架,结合ASR、NER、LLM-RAG和说话人嵌入,通过规则风险评分检测保险欺诈中的叙述复用、结构不一致和跨案件语音重复。

Comments 10 pages, 8 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05818 2026-04-15 cs.CV cs.CL cs.IR 82%

WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering

WikiSeeker: 重新思考视觉语言模型在基于知识的视觉问答中的作用

Yingjian Zhu, Xinming Wang, Kun Ding, Ying Wang, Bin Fan, Shiming Xiang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室 (MAIS))

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.IR、cs.CL

AI总结 本文提出WikiSeeker框架,通过引入多模态检索器和重新定义视觉语言模型的角色,提升多模态检索性能和答案质量,实现在EVQA、InfoSeek和M2KR数据集上的最优表现。

Comments Accepted by ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03967 2026-03-05 cs.CV 82%

UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization

UniRain: 基于RAG的数据蒸馏和多目标重加权优化的统一图像去雨

Qianfeng Yang, Qiyuan Guan, Xiang Chen, Jiyu Jin, Guiyue Jin, Jiangxin Dong

机构 * Dalian Polytechnic University(大连理工大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(abstract)

AI总结 UniRain通过基于RAG的数据蒸馏和多目标重加权优化,实现统一图像去雨,有效应对不同雨损场景。

Comments Accepted by CVPR 2026; Project Page: https://github.com/QianfengY/UniRain

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00511 2026-03-03 cs.CV cs.LG 82%

Multimodal Adaptive Retrieval Augmented Generation through Internal Representation Learning

多模态自适应检索增强生成通过内部表示学习

Ruoshuang Du, Xin Sun, Qiang Liu, Bowen Song, Zhongqi Chen, Weiqiang Wang, Liang Wang

机构 * Shanghaitech University School of Information Science(上海科技大学信息科学学院) Chinese Academy of Sciences Institute of Automation(中国科学院自动化研究所)

专题命中 多模态RAG :retrieval augmented generation(title,abstract);RAG(abstract)

AI总结 本文提出MMA-RAG,通过动态评估模型内部知识置信度,提升视觉问答系统在多模态场景下的检索增强生成性能。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15650 2026-02-18 cs.CV 82%

Concept-Enhanced Multimodal RAG: Towards Interpretable and Accurate Radiology Report Generation

概念增强的多模态RAG:迈向可解释且准确的放射科报告生成

Marco Salmè, Federico Siciliano, Fabrizio Silvestri, Paolo Soda, Rosa Sicilia, Valerio Guarrasi

机构 * Department of Engineering(工程系) Research Unit of Artificial Intelligence and Computer Systems(人工智能与计算机系统研究单位) Università Campus Bio-Medico of Roma(罗马大学生物医学校园) Department of Computer, Control and Management Engineering(计算机、控制与管理工程系) Sapienza University of Rome(罗马萨皮恩扎大学) Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering(诊断与介入、辐射物理、生物医学工程系) Umeå University(乌梅拉大学) UniCamillus-Saint Camillus International University of Health Sciences(UniCamillus-圣卡米卢斯国际健康科学大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract)

AI总结 概念增强的多模态RAG通过分解视觉表示为可解释的临床概念,提升放射科报告生成的可解释性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00030 2026-02-10 cs.LG 82%

RAPTOR-AI for Disaster OODA Loop: Hierarchical Multimodal RAG with Experience-Driven Agentic Decision-Making

RAPTOR-AI用于灾难OODA循环:基于经验驱动的多模态RAG框架

Takato Yasuno

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract)

AI总结 RAPTOR-AI通过分层多模态RAG框架,结合经验驱动的代理控制和LoRA适应,提升灾害响应中的检索精度、情境基础性和任务分解准确性。

Comments 8 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02371 2025-11-05 cs.LG 82%

LUMA-RAG: Lifelong Multimodal Agents with Provably Stable Streaming Alignment

Rohan Wandre, Yash Gajewar, Namrata Patel, Vivek Dhalkari

机构 * Dept. of Computer Engineering(计算机工程系) SIES Graduate School of Technology(SIES技术研究生学院) Bharatiya Vidya Bhavan's Sardar Patel Institute of Technology(巴哈里亚·维达·巴万学院萨达尔·帕特尔技术学院)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01546 2025-08-05 cs.CV 82%

E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation

Zeyu Xu, Junkang Zhang, Qiang Wang, Yi Liu

机构 * Zeyu Xu(作者) Junkang Zhang(作者) Qiang Wang(作者) Yi Liu(作者)

专题命中 多模态RAG :retrieval augmented generation(title);retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07902 2025-07-11 cs.CV 82%

MIRA: A Novel Framework for Fusing Modalities in Medical RAG

Jinhong Wang, Tajamul Ashraf, Zongyan Han, Jorma Laaksonen, Rao Mohammad Anwer

机构 * Department of Computer Vision, MBZUAI(视觉计算系,MBZUAI) Department of Computer Science, Aalto University(计算机科学系,阿alto大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10320 2025-04-15 cs.CV 82%

SlowFastVAD: Video Anomaly Detection via Integrating Simple Detector and RAG-Enhanced Vision-Language Model

Zongcan Ding, Haodong Zhang, Peng Wu, Guansong Pang, Zhiwei Yang, Peng Wang, Yanning Zhang

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06254 2025-03-17 cs.CR cs.LG 82%

Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation

Yinuo Liu, Zenghui Yuan, Guiyao Tie, Jiawen Shi, Pan Zhou, Lichao Sun, Neil Zhenqiang Gong

专题命中 多模态RAG :retrieval augmented generation(title);retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02800 2025-03-12 cs.LG cs.CE 82%

RAAD-LLM: Adaptive Anomaly Detection Using LLMs and RAG Integration

Alicia Russell-Gilbert, Sudip Mittal, Shahram Rahimi, Maria Seale, Joseph Jabour, Thomas Arnold, Joshua Church

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract)

Comments arXiv admin note: substantial text overlap with arXiv:2411.00914

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20927 2024-12-31 cs.CV 82%

Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering

Junxiao Xue, Quan Deng, Fei Yu, Yanhao Wang, Jun Wang, Yuehua Li

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract)

Comments 6 pages, 3 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03714 2024-06-07 cs.SD eess.AS 82%

Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining

Jinlong Xue, Yayue Deng, Yingming Gao, Ya Li

专题命中 多模态RAG :retrieval augmented generation(title,abstract);RAG(abstract)

Comments Accepted by Interspeech 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04625 2026-08-06 cs.AI 新提交 81%

A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing

A/B Agent:面向工业A/B测试中策略迭代的自进化智能体

Zhuohang Jiang, Yuxin Chen, Yongsen Pan, Zheng Hu, Wenqi Fan, Qing Li, Hongyang Wang, Jun Wang, Wenwu Ou

机构 * The Hong Kong Polytechnic University(香港理工大学) Kuaishou Technology(快手科技) University of Electronic Science and Technology of China(电子科技大学) Southwest Jiaotong University(西南交通大学)

专题命中 多模态RAG :RAG(summary_cn,abstract);分类 cs.AI

AI总结 该研究提出A/B Agent智能体,通过层级经验树与Tree-RAG技术优化工业A/B测试的策略迭代,在短视频电商场景实现GMV提升4.829%且护栏指标正向增益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24554 2026-07-28 cs.IR cs.CV 新提交 81%

DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding

DeCoRAG:用于复杂文档理解的认知解耦和语义感知裁剪

Shuo Wang, Kai Zhang, Wenyuan Huang, Yizheng Yu, Xia Liao, Junming Su, Qing Wang, Fang Xi

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);hybrid retrieval(abstract);分类 cs.IR

AI总结 研究针对复杂文档理解中多模态检索增强生成的难题,提出DeCoRAG方法,通过认知解耦、建立语义锚及区域感知裁剪机制,提升语义通过率,减少提示令牌,在复杂文档基准测试中取得良好效果。

Comments 11 pages, 4 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02185 2026-07-03 cs.CV cs.AI 新提交 81%

RadiomicNet: A Hybrid Radiomics-Guided Lightweight Architecture for Interpretable Medical Image Segmentation

RadiomicNet: 一种混合放射组学引导的轻量级可解释医学图像分割架构

Mohammad Amanour Rahman

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Ahsanullah University of Science and Technology (AUST)(阿萨努拉科学与技术大学)

专题命中 多模态RAG :RAG(summary_cn,abstract);分类 cs.AI

AI总结 提出RadiomicNet,通过放射组学注意力门(RAG)将手工放射组学特征集成到轻量级编码器-解码器中,实现可解释分割,在BUSI和Kvasir-SEG数据集上分别达到0.763和0.854的Dice系数,参数仅3.27M。

Comments Accepted at the IEEE ICIP 2026 LBDL 2 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25496 2026-06-25 cs.IR 新提交 81%

Recommendation as Generation: Unifying Personalized Video Generation and Recommendation at Industrial Scale

推荐即生成:在工业规模上统一个性化视频生成与推荐

Yanhua Cheng, Bo Wang, Haotian Zhang, Xinyuan Gao, Zhihui Yin, Ben Xue, Yongzhi Li, Jieting Xue, Ye Ma, Minquan Wang, Jiahui Li, Tianyu Xu, Zhiqiang Liu, Xiao Lin, Shiyang Wen, Changcheng Li, Liu Liu, Quan Chen, Peng Jiang, Kun Gai

专题命中 多模态RAG :RAG(summary_cn,abstract);分类 cs.IR

AI总结 提出推荐即生成(RaG)范式,通过共享语义ID统一生成式推荐与视频生成,实现个性化视频按需生成,并在工业平台部署,广告收入提升1.87%。

Comments Project page: https://recommendation-as-generation.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07252 2026-06-09 cs.IR 新提交 81%

Constrained Dominant Sets for Multimodal Document Question Answering

约束主导集用于多模态文档问答

Ambuj Mehrish, Sebastiano Vascon

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.IR

AI总结 提出基于查询增强亲和图的约束主导集检索方法,通过谱边界自动平衡相关性与冗余性,利用复制动态实现全局均衡,无需训练,在多模态文档问答中取得新最优结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27378 2026-05-28 cs.CL cs.CV cs.MA 81%

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

OralAgent: 融合推理、工具与知识的交互式牙科影像分析

Jing Hao, Siyuan Dai, Yongxin Zhang, Yuci Liang, Jiamin Wu, Jiahao Bao, Yuxuan Fan, Zanting Ye, Yanpeng Sun, Xinyu Zhang, Ming Hu, Liang Zhan, James Kit Hon Tsoi, Linlin Shen, Junjun He, Kuo Feng Hung

机构 * Faculty of Dentistry, the University of Hongkong, Hong Kong SAR, China(香港大学牙科学院,中国香港特别行政区) Department of Electrical and Computer Engineering, University of Pittsburgh, Pittsburgh, PA, USA(匹兹堡大学电气与计算机工程系,美国宾夕法尼亚州匹兹堡) Shenzhen University, China(深圳大学,中国) Department of Craniomaxillofacial Surgery, Shanghai Ninth People’s Hospital, China(上海第九人民医院口腔颌面外科部,中国) Nanyang technological University, Singapore(南洋理工大学,新加坡) School of Biomedical Engineering, Southern Medical University, China(南方医科大学生物医学工程学院,中国) Singapore University of Technology and Design, Singapore(新加坡科技设计大学,新加坡) University of Auckland, new zealand(奥克兰大学,新西兰) Shanghai Artificial Intelligence Laboratory , China(上海人工智能实验室,中国)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);knowledge retrieval(abstract);分类 cs.CL

AI总结 提出首个牙科专用AI智能体OralAgent,通过集成22种视觉分析工具和368本经典牙科教科书,实现多模态推理、工具决策与知识检索的自动化框架,在多个基准上达到最优性能。

Comments 14 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏