arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 546 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 546 篇

2507.05461 2025-07-23 cs.HC 67%

GLOSS: Group of LLMs for Open-Ended Sensemaking of Passive Sensing Data for Health and Wellbeing

Akshat Choube, Ha Le, Jiachen Li, Kaixin Ji, Vedant Das Swain, Varun Mishra

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15266 2025-07-22 cs.RO cs.SY eess.SY 67%

VLM-UDMC: VLM-Enhanced Unified Decision-Making and Motion Control for Urban Autonomous Driving

Haichao Liu, Haoren Guo, Pei Liu, Benshan Ma, Yuxiang Zhang, Jun Ma, Tong Heng Lee

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

Comments 14 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09149 2025-06-23 cs.CV cs.MM 67%

Memory-enhanced Retrieval Augmentation for Long Video Understanding

Huaying Yuan, Zheng Liu, Minghao Qin, Hongjin Qian, Yan Shu, Zhicheng Dou, Ji-Rong Wen, Nicu Sebe

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学光华学院人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) University of Trento(特伦托大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06399 2025-05-13 cs.RO 67%

LLM-Land: Large Language Models for Context-Aware Drone Landing

Siwei Cai, Yuwei Wu, Lifeng Zhou

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Drexel University(德雷塞尔大学) Department of Electrical and Systems Engineering(电气与系统工程系) University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03556 2025-05-07 cs.IT math.IT 67%

A Comprehensive Survey of Large AI Models for Future Communications: Foundations, Applications and Challenges

Feibo Jiang, Cunhua Pan, Li Dong, Kezhi Wang, Merouane Debbah, Dusit Niyato, Zhu Han

专题命中 多模态RAG :retrieval augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07110 2025-04-22 eess.SY cs.SY 67%

Bio-Eng-LMM AI Assist chatbot: A Comprehensive Tool for Research and Education

Ali Forootani, Danial Esmaeili Aliabadi, Daniela Thraen

专题命中 多模态RAG :retrieval augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13190 2025-04-21 cs.NI eess.SP 67%

Cellular-X: An LLM-empowered Cellular Agent for Efficient Base Station Operations

Liujianfu Wang, Xinyi Long, Yuyang Du, Xiaoyan Liu, Kexin Chen, Soung Chang Liew

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

Comments MobiSys ’25, June 23-27, 2025, Anaheim, CA, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13964 2025-03-19 cs.LG 67%

MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding

Siwei Han, Peng Xia, Ruiyi Zhang, Tong Sun, Yun Li, Hongtu Zhu, Huaxiu Yao

专题命中 多模态RAG :retrieval augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13508 2025-02-24 cs.RO 67%

VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation

Wei Zhao, Pengxiang Ding, Min Zhang, Zhefei Gong, Shuanghao Bai, Han Zhao, Donglin Wang

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

Comments Accepted as a conference paper at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11024 2025-02-18 cs.CV 67%

TPCap: Unlocking Zero-Shot Image Captioning with Trigger-Augmented and Multi-Modal Purification Modules

Ruoyu Zhang, Lulu Wang, Yi He, Tongling Pan, Zhengtao Yu, Yingna Li

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06919 2025-01-14 cs.RO 67%

Shake-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Manipulations and Liquid Mixing

Muhamamd Haris Khan, Selamawit Asfaw, Dmitrii Iarchuk, Miguel Altamirano Cabrera, Luis Moreno, Issatay Tokmurziyev, Dzmitry Tsetserukou

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

Comments Accepted to IEEE/ACM HRI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00847 2024-12-31 cs.DB cs.AI cs.IR 67%

The Design of an LLM-powered Unstructured Analytics System

Eric Anderson, Jonathan Fritz, Austin Lee, Bohou Li, Mark Lindblad, Henry Lindeman, Alex Meyer, Parth Parmar, Tanvi Ranade, Mehul A. Shah, Benjamin Sowell, Dan Tecuci, Vinayak Thapliyal, Matt Welsh

专题命中 多模态RAG :RAG(abstract);分类 cs.IR、cs.AI、cs.DB

Comments Included in the proceedings of The Conference on Innovative Data Systems Research (CIDR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12735 2024-12-03 cs.CV 67%

EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Yibin Yan, Weidi Xie

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

Comments Accepted by EMNLP 2024 findings; Project Page: https://go2heart.github.io/echosight

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14083 2024-09-24 cs.CV 67%

SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information

Jiashuo Sun, Jihai Zhang, Yucheng Zhou, Zhaochen Su, Xiaoye Qu, Yu Cheng

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

Comments 19 pages, 9 tables, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14944 2024-09-11 cs.CV 67%

Automatic Generation of Fashion Images using Prompting in Generative Machine Learning Models

Georgia Argyrou, Angeliki Dimitriou, Maria Lymperaiou, Giorgos Filandrianos, Giorgos Stamou

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

Journal ref ECCVW 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10153 2024-07-16 cs.CV cs.LG 67%

Improving Medical Multi-modal Contrastive Learning with Expert Annotations

Yogesh Kumar, Pekka Marttinen

专题命中 多模态RAG :retrieval augmented generation(abstract);RAG(abstract)

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18039 2024-06-27 physics.med-ph 67%

Diagnosis Assistant for Liver Cancer Utilizing a Large Language Model with Three Types of Knowledge

Xuzhou Wu, Guangxin Li, Xing Wang, Zeyu Xu, Yingni Wang, Jianming Xian, Xueyu Wang, Gong Li, Kehong Yuan

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20834 2024-06-03 cs.CV 67%

Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning

Cheng Tan, Jingxuan Wei, Linzhuang Sun, Zhangyang Gao, Siyuan Li, Bihui Yu, Ruifeng Guo, Stan Z. Li

专题命中 多模态RAG :retrieval-augmented generation(abstract);RAG(abstract)

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17136 2023-11-30 cs.CV cs.AI cs.CL cs.IR 67%

UniIR: Training and Benchmarking Universal Multimodal Information Retrievers

Cong Wei, Yang Chen, Haonan Chen, Hexiang Hu, Ge Zhang, Jie Fu, Alan Ritter, Wenhu Chen

专题命中 多模态RAG :retriever(abstract);分类 cs.IR、cs.CL、cs.AI

Comments Our code and dataset are available on this project page: https://tiger-ai-lab.github.io/UniIR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04240 2026-06-04 cs.CV cs.AI cs.CL 62%

Overview of the EReL@MIR 2025 Multimodal Document Retrieval Challenge (Track 1)

EReL@MIR 2025 多模态文档检索挑战赛(赛道1)概述

Jingbiao Mei

机构 * University of Cambridge(剑桥大学) Cambridge United Kingdom(剑桥英国)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文介绍了EReL@MIR 2025多模态文档检索挑战赛(赛道1)的设计、数据集、评估协议、最终排名及前三名获胜系统的分析,所有系统均基于Qwen2-VL系列解码器多模态大语言模型嵌入器。

Comments MDR Challenge Report at WWW2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13719 2026-03-25 cs.CV cs.AI cs.IR 62%

Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search

基于音频视觉实体凝聚力与代理搜索的层次化长视频理解

Xinlei Yin, Xiulian Peng, Xiao Li, Zhiwei Xiong, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文提出HAVEN框架,通过整合音频视觉实体凝聚力与层次化视频索引与代理搜索,解决长视频理解中的信息碎片化和全局一致性问题,实验显示在LVBench上达到84.1%的准确率。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15623 2026-03-18 cs.IR cs.AI 62%

Finder: A Multimodal AI-Powered Search Framework for Pharmaceutical Data Retrieval

Finder:一种多模态AI驱动的制药数据检索框架

Suyash Mishra, Srikanth Patil, Satyanarayan Pati, Sagar Sahu, Baddu Narendra

机构 * Researcher, Global Product Strategy F. Hoffmann-La Roche Ltd. Basel, Switzerland(全球产品战略研究员 罗氏有限公司 巴塞尔,瑞士) Associate Vice President (Gen AI) Involead Services Pvt Ltd. Pune, India(高级副总裁(生成式人工智能) Involead服务私人有限公司 普纳,印度) Lead Data Scientist Involead Services Pvt Ltd. Delhi, India(首席数据科学家 Involead服务私人有限公司 德里,印度) Data Scientist Involead Services Pvt Ltd. Bhubaneswar, India(数据科学家 Involead服务私人有限公司 奇塔拉,印度)

专题命中 多模态RAG :vector search(abstract);分类 cs.IR、cs.AI

AI总结 Finder利用混合向量搜索统一文本、图像、音频和视频的检索,通过稀疏词汇和密集语义模型提升多模态内容处理能力,支持自然语言推理搜索,已处理超过29万文档、3.1万视频和1192音频文件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02454 2025-12-16 cs.CL cs.AI 62%

Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework

多模态深度研究员:从零开始生成文本-图表交织报告的代理框架

Zhaorui Yang, Bo Pan, Han Wang, Yiyao Wang, Xingyu Liu, Luoxuan Weng, Yingchaojie Feng, Haozhe Feng, Minfeng Zhu, Bo Zhang, Wei Chen

专题命中 多模态RAG :retrieval augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 多模态深度研究员通过代理框架实现从零开始生成文本-图表交织报告,利用FDV结构化文本表示提升可视化生成质量。

Comments AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20513 2025-11-26 cs.CV cs.AI cs.CL cs.HC 62%

DesignPref: Capturing Personal Preferences in Visual Design Generation

DesignPref: 在视觉设计生成中捕捉个人偏好

Yi-Hao Peng, Jeffrey P. Bigham, Jason Wu

机构 * Carnegie Mellon University(卡内基梅隆大学) Apple(苹果公司)

专题命中 多模态RAG :RAG(abstract);分类 cs.CL、cs.AI

AI总结 DesignPref通过引入包含20名专业设计师多级评分的12,000对UI设计比较数据集,研究个性化视觉设计评估,展示个性化模型在预测个体偏好上的优越性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15370 2025-11-20 cs.CL cs.AI 62%

The Empowerment of Science of Science by Large Language Models: New Tools and Methods

Guoqiang Liang, Jingqian Gong, Mengxuan Li, Gege Lin, Shuo Zhang

专题命中 多模态RAG :retrieval augmented generation(abstract);分类 cs.CL、cs.AI

Comments The manuscript is currently ongoing the underreview process of the journal of information science

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16470 2025-11-10 cs.IR cs.CL cs.CV 62%

Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering

Kuicai Dong, Yujing Chang, Shijie Huang, Yasheng Wang, Ruiming Tang, Yong Liu

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

Comments Paper accepted to NeurIPS 2025 DB

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06888 2025-10-09 cs.IR cs.AI 62%

M3Retrieve: Benchmarking Multimodal Retrieval for Medicine

Arkadeep Acharya, Akash Ghosh, Pradeepika Verma, Kitsuchart Pasupa, Sriparna Saha, Priti Singh

机构 * Indian Institute of Technology Patna(印度理工学院帕纳瓦分校) King Mongkut’s Institute of Technology Ladkrabang(拉差丹awan技术大学)

专题命中 多模态RAG :RAG(abstract);分类 cs.IR、cs.AI

Comments EMNLP Mains 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08897 2025-09-12 cs.CV cs.AI cs.CL cs.MM 62%

Recurrence Meets Transformers for Universal Multimodal Retrieval

Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

机构 * Department of Education and Humanities, University of Modena and Reggio Emilia(教育与人文学院, Modena and Reggio Emilia大学)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21604 2025-06-30 cs.IR cs.AI cs.CV cs.HC cs.LG 62%

Evaluating VisualRAG: Quantifying Cross-Modal Performance in Enterprise Document Understanding

Varun Mannam, Fang Wang, Xin Chen

机构 * Amazon(亚马逊)

专题命中 多模态RAG :RAG(abstract);分类 cs.IR、cs.AI

Comments Conference: KDD conference workshop: https://kdd-eval-workshop.github.io/genai-evaluation-kdd2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20490 2025-06-13 cs.CV cs.AI cs.CL 62%

EgoNormia: Benchmarking Physical Social Norm Understanding

MohammadHossein Rezaei, Yicheng Fu, Phil Cuvin, Caleb Ziems, Yanzhe Zhang, Hao Zhu, Diyi Yang

专题命中 多模态RAG :RAG(abstract);分类 cs.CL、cs.AI

Comments V4, fixes to title and formatting

详情

展开后加载摘要…

URL PDF HTML 收藏