arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 546 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 546 篇

2508.12263 2025-09-01 cs.CV cs.AI 57%

Region-Level Context-Aware Multimodal Understanding

Hongliang Wei, Xianqi Zhang, Xingtao Wang, Xiaopeng Fan, Debin Zhao

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院长) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19319 2025-08-28 eess.IV cs.AI cs.CV 57%

MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction

Pardis Moradbeiki, Nasser Ghadiri, Sayed Jalal Zahabi, Uffe Kock Wiil, Kristoffer Kittelmann Brockhattingen, Ali Ebrahimi

机构 * Department of Electrical and Computer Engineering, Isfahan University of Technology(电气与计算机工程系,伊斯法罕技术大学) SDU Health Informatics and Technology, The Maersk Mc-Kinney Moller Institute, University of Southern Denmark(南部丹麦大学健康信息学与技术,马士基麦金尼莫勒研究所) Geriatric Research Unit, Department of Clinical Research, University of Southern Denmark(老年医学研究单元,临床研究系,南部丹麦大学)

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18108 2025-08-26 cs.CL 57%

SentiMM: A Multimodal Multi-Agent Framework for Sentiment Analysis in Social Media

Xilai Xu, Zilin Zhao, Chengye Song, Zining Wang, Jinhe Qiang, Jiongrui Yan, Yuhuai Lin

机构 * College of Information and Electrical Engineering, China Agricultural University(信息与电气工程学院,中国农业大学) College of Software, Jilin University(软件学院,吉林大学) College of Communication Engineering, Jilin University(通信工程学院,吉林大学)

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06328 2025-08-11 cs.IR 57%

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation

Zhiyou Xiao, Qinhan Yu, Binghui Li, Geng Chen, Chong Chen, Wentao Zhang

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12884 2025-07-01 cs.LG cs.AI cs.CV 57%

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks

Yuanze Hu, Zhaoxin Fan, Xinyu Wang, Gen Li, Ye Qiu, Zhichao Yang, Wenjun Wu, Kejian Wu, Yifan Sun, Xiaotie Deng, Jin Dong

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算先进创新中心) Beihang University(北京航空航天大学) Hangzhou International Innovation Institute(杭州国际创新研究院) Xreal Renmin University(中国人民大学) Peking University(北京大学) Beijing Academy of Blockchain and Edge Computing (BABEC)(北京区块链与边缘计算研究院)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12831 2025-07-01 eess.IV cs.AI cs.CV 57%

Segment as You Wish -- Free-Form Language-Based Segmentation for Medical Images

Longchao Da, Rui Wang, Xiaojian Xu, Parminder Bhatia, Taha Kass-Hout, Hua Wei, Cao Xiao

机构 * Arizona State University(亚利桑那州立大学)

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 19 pages, 9 as main content. The paper was accepted to KDD2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10756 2025-06-13 cs.RO cs.AI 57%

Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding

Yuhang Zhang, Haosheng Yu, Jiaping Xiao, Mir Feroskhan

机构 * School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学)

专题命中 多模态RAG :retriever(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10913 2025-06-11 cs.SD cs.AI eess.AS 57%

Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning

Choi Changin, Lim Sungjun, Rhee Wonjong

机构 * Interdisciplinary Program in Artificial Intelligence(人工智能交叉学科项目) Department of Intelligence and Information(智能与信息系)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02470 2025-06-04 cs.AI 57%

A Smart Multimodal Healthcare Copilot with Powerful LLM Reasoning

Xuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan Miao

机构 * Joint NTU-UBC Research Centre of Excellence in Active Living for the Elderly (LILY), NTU(联合NTU-UBC老龄化积极生活卓越研究中心(LILY),NTU) College of Computing and Data Science, Nanyang Technological University (NTU), Singapore(计算与数据科学学院,南洋理工大学(NTU),新加坡) Tan Tock Seng Hospital, Singapore(坦 tok sing 医院,新加坡) Woodlands Health, Singapore(伍德兰兹健康,新加坡)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14318 2025-06-03 cs.CV cs.CL 57%

RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection

Wenjun Hou, Yi Cheng, Kaishuai Xu, Heng Li, Yan Hu, Wenjie Li, Jiang Liu

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Research Institute of Trustworthy Autonomous Systems and Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学可信自主系统研究院和计算机科学与工程系) School of Computer Science, University of Nottingham Ningbo China(宁波大学计算机学院)

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

Comments Accepted to ACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02466 2025-05-06 cs.IR 57%

Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

Xueguang Ma, Luyu Gao, Shengyao Zhuang, Jiaqi Samantha Zhan, Jamie Callan, Jimmy Lin

专题命中 多模态RAG :retriever(abstract);分类 cs.IR

Comments Accepted in SIGIR 2025 (Demo)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16723 2025-04-24 cs.CV cs.AI 57%

Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering

Ali Anaissi, Junaid Akram, Kunal Chaturvedi, Ali Braytee

机构 * The University of Sydney, School of Computer Science(悉尼大学计算机科学学院) University of Technology Sydney, School of Computer Science(新南威尔士大学技术学院) University of Technology Sydney, TD School(新南威尔士大学TD学院) Australian Catholic University, Peter Faber Business School(澳大利亚天主教大学彼得·法伯商学院)

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 13 pages, 2 figures, 2025 International Conference on Computational Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00750 2025-04-17 cs.AI 57%

Beyond Text: Implementing Multimodal Large Language Model-Powered Multi-Agent Systems Using a No-Code Platform

Cheonsu Jeong

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 22 pages, 27 figures

Journal ref 2025 Journal of Intelligence and Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03241 2025-04-07 cs.CV cs.AI cs.LG 57%

Rotation Invariance in Floor Plan Digitization using Zernike Moments

Marius Graumann, Jan Marius Stürmer, Tobias Koch

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12836 2025-03-05 cs.CV cs.AI cs.LG 57%

AI-based association analysis for medical imaging using latent-space geometric confounder correction

Xianjing Liu, Bo Li, Meike W. Vernooij, Eppo B. Wolvius, Gennady V. Roshchupkin, Esther E. Bron

专题命中 多模态RAG :vector search(abstract);分类 cs.AI

Comments Accepted by Medical Image Analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05846 2025-03-04 cs.CV cs.CL 57%

Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines

Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad, Yonatan Belinkov

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

Comments Published in: ACL 2024 Project webpage: tokeron.github.io/DiffusionLensWeb

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08250 2025-02-24 cs.HC cs.AI 57%

OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering

Jiahao Nick Li, Zhuohao Jerry Zhang, Jiaju Ma

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments Paper accepted to the 2025 CHI Conference on Human Factors in Computing Systems (CHI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07365 2025-02-18 cs.IR cs.LG 57%

Multimodal semantic retrieval for product search

Dong Liu, Esther Lopez Ramos

专题命中 多模态RAG :dense retrieval(abstract);分类 cs.IR

Comments Accepted at EReL@MIR WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06781 2025-01-27 cs.AI 57%

Eliza: A Web3 friendly AI Agent Operating System

Shaw Walters, Sam Gao, Shakker Nerd, Feng Da, Warren Williams, Ting-Chien Meng, Amie Chow, Hunter Han, Frank He, Allen Zhang, Ming Wu, Timothy Shen, Maxwell Hu, Jerry Yan

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00846 2024-12-03 cs.AI 57%

Improving Multimodal LLMs Ability In Geometry Problem Solving, Reasoning, And Multistep Scoring

Avinash Anand, Raj Jaiswal, Abhishek Dharmadhikari, Atharva Marathe, Harsh Parimal Popat, Harshil Mital, Kritarth Prasad, Rajiv Ratn Shah, Roger Zimmermann

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16592 2024-10-23 cs.LG cs.CL cs.CY 57%

ViMGuard: A Novel Multi-Modal System for Video Misinformation Guarding

Andrew Kan, Christopher Kan, Zaid Nabulsi

专题命中 多模态RAG :retrieval augmented generation(abstract);分类 cs.CL

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13510 2024-10-18 cs.CL cs.CV 57%

GeoCoder: Solving Geometry Problems by Generating Modular Code through Vision-Language Models

Aditya Sharma, Aman Dalmia, Mehran Kazemi, Amal Zouaq, Christopher J. Pal

专题命中 多模态RAG :RAG(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18202 2024-10-01 cs.AI cs.MM 57%

WorldGPT: Empowering LLM as Multimodal World Model

Zhiqi Ge, Hongzhe Huang, Mingze Zhou, Juncheng Li, Guoming Wang, Siliang Tang, Yueting Zhuang

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.AI

Comments update v2

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11281 2024-09-18 cs.IR 57%

Beyond Relevance: Improving User Engagement by Personalization for Short-Video Search

Wentian Bao, Hu Liu, Kai Zheng, Chao Zhang, Shunyu Zhang, Enyun Yu, Wenwu Ou, Yang Song

专题命中 多模态RAG :dense retrieval(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06450 2024-09-11 cs.RO cs.AI cs.ET 57%

Multimodal Large Language Model Driven Scenario Testing for Autonomous Vehicles

Qiujing Lu, Xuanhan Wang, Yiwei Jiang, Guangming Zhao, Mingyue Ma, Shuo Feng

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14723 2024-08-28 cs.CV cs.IR 57%

Snap and Diagnose: An Advanced Multimodal Retrieval System for Identifying Plant Diseases in the Wild

Tianqi Wei, Zhi Chen, Xin Yu

专题命中 多模态RAG :retriever(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09272 2024-07-26 cs.CV cs.AI cs.SD eess.AS 57%

Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos

Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue, Wei-Ning Hsu, David Harwath, Kristen Grauman

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

Comments Project page: https://vision.cs.utexas.edu/projects/action2sound. ECCV 2024 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15218 2024-02-26 cs.CR cs.CL cs.CV 57%

BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators

Yu Tian, Xiao Yang, Yinpeng Dong, Heming Yang, Hang Su, Jun Zhu

专题命中 多模态RAG :retriever(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19787 2024-01-08 cs.CV cs.AI 57%

DeepMerge: Deep-Learning-Based Region-Merging for Image Segmentation

Xianwei Lv, Claudio Persello, Wangbin Li, Xiao Huang, Dongping Ming, Alfred Stein

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13820 2023-08-29 cs.IR 57%

Video and Audio are Images: A Cross-Modal Mixer for Original Data on Video-Audio Retrieval

Zichen Yuan, Qi Shen, Bingyi Zheng, Yuting Liu, Linying Jiang, Guibing Guo

专题命中 多模态RAG :retriever(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏