arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 546 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 546 篇

2601.13801 2026-01-21 cs.RO 50%

HoverAI: An Embodied Aerial Agent for Natural Human-Drone Interaction

HoverAI: 一种用于自然人-无人机交互的具身空中代理

Yuhua Jin, Nikita Kuzmin, Georgii Demianchuk, Mariya Lezina, Fawad Mehboob, Issatay Tokmurziyev, Miguel Altamirano Cabrera, Muhammad Ahsan Mustafa, Dzmitry Tsetserukou

机构 * Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Skolkovo Institute of Science and Technology(斯克尔科沃信息科技研究所)

专题命中 多模态RAG :RAG(abstract)

AI总结 HoverAI通过结合无人机移动、视觉投影和对话式AI,实现了人-无人机自然交互的具身代理,提升了空间感知与社交响应能力。

Comments This paper has been accepted for publication at LBR HRI 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11537 2026-01-21 cs.HC 50%

Building AI-based advisory services for smallholder farmers: Technical learnings from the AIEP Initiative

为小农户建设基于AI的咨询服务:AIEP计划的技术经验

Stewart Collis, Florence Kinyua, Vikram Kumar, Howard Lakougna, Christian Merz, Kirti Pandey, Christian Resch

专题命中 多模态RAG :RAG(abstract)

AI总结 AIEP计划通过AI技术为小农户提供农业咨询服务,发现多语言语音交互和语料库编纂是关键挑战,同时强调数据共享和评估基准的重要性。

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23044 2026-01-19 cs.CV 50%

Video-Browser: Towards Agentic Open-web Video Browsing

Video-Browser: 向具有代理能力的开放网页视频浏览迈进

Zhengyang Liang, Yan Shu, Xiangrui Liu, Minghao Qin, Kaixin Liang, Nicu Sebe, Zheng Liu, Lizi Liao

专题命中 多模态RAG :RAG(abstract)

AI总结 Video-Browser通过金字塔感知技术,在开放式网络视频浏览中实现37.5%的相对提升,同时减少58.3%的token消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04540 2025-12-17 cs.CV 50%

VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management

VideoMem: 通过自适应内存管理增强超长视频理解

Hongbo Jin, Qingyuan Wang, Wenhao Zhang, Yang Liu, Sijie Cheng

机构 * School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学) Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学)

专题命中 多模态RAG :RAG(abstract)

AI总结 VideoMem通过自适应内存管理框架,有效提升超长视频理解任务的性能,采用PRPO算法和两个核心模块实现高效训练和长期记忆保留。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02395 2025-12-09 cs.CV 50%

Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch

Skywork-R1V4:通过图像与深度研究交织思考实现代理多模态智能

Yifan Zhang, Liang Hu, Haofeng Sun, Peiyu Wang, Yichen Wei, Shukang Yin, Jiangbo Pei, Wei Shen, Peng Xia, Yi Peng, Tianyidan Xie, Eric Li, Yang Liu, Xuchen Song, Yahui Zhou

机构 * Skywork AI

专题命中 多模态RAG :knowledge retrieval(abstract)

AI总结 Skywork-R1V4通过交织推理实现多模态代理智能,仅用监督学习在少数据上训练,超越现有模型在多个基准测试中的表现。

Comments 21 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18883 2025-11-24 cs.CV 50%

Universal Video Temporal Grounding with Generative Multi-modal Large Language Models

通用视频时间定位与生成多模态大语言模型

Zeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang, Yanfeng Wang, Weidi Xie

机构 * SAI, Shanghai Jiao Tong University(上海交通大学SAI实验室) ByteDance Seed(字节跳动种子)

专题命中 多模态RAG :retriever(abstract)

AI总结 本文提出UniTime模型,利用生成多模态大语言模型实现通用视频时间定位,有效处理多类型视频并提升VideoQA任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06752 2025-11-11 cs.CV 50%

Med-SORA: Symptom to Organ Reasoning in Abdomen CT Images

You-Kyoung Na, Yeong-Jun Cho

机构 * Chonnam National University(全南国立大学)

专题命中 多模态RAG :RAG(abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05199 2025-11-10 cs.RO 50%

Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation

Yichen Zhu, Feifei Feng

机构 * Midea Group, AI Research Center(美的集团人工智能研究中心)

专题命中 多模态RAG :retriever(abstract)

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26386 2025-10-29 cs.CV 50%

PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer

Zhiwei Yang, Chen Gao, Mike Zheng Shou

机构 * Xidian University(西安电子科技大学) Show Lab, National University of Singapore(新加坡国立大学Show实验室)

专题命中 多模态RAG :RAG(abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16510 2025-10-21 q-bio.BM 50%

CryoDyna: Multiscale end-to-end modeling of cryo-EM macromolecule dynamics with physics-aware neural network

Chengwei Zhang, Shimian Li, Yihao Niu, Zhen Zhu, Sihao Yuan, Sirui Liu, Yi Qin Gao

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19203 2025-09-24 cs.CV 50%

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

Ioanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

机构 * Queen Mary University of London(伦敦女王大学) Samsung AI Centre(三星人工智能中心) Technical University of Iași(伊阿苏技术大学)

专题命中 多模态RAG :retriever(abstract)

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01960 2025-09-23 cs.LG 50%

MPIC: Position-Independent Multimodal Context Caching System for Efficient MLLM Serving

Shiju Zhao, Junhao Hu, Rongxiao Huang, Jiaqi Zheng, Guihai Chen

机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) Nanjing University(南京大学) School of Computer Science(计算机学院)

专题命中 多模态RAG :retrieval-augmented generation(abstract)

Comments 17 pages, 13 figures, the second version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13899 2025-09-18 cs.HC 50%

AI as a teaching tool and learning partner

Steven Watterson, Sarah Atkinson, Elaine Murray, Andrew McDowell

专题命中 多模态RAG :RAG(abstract)

Comments 6 Pages, 1 Figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01970 2025-08-29 cs.LG 50%

Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling

Rituparna Datta, Jiaming Cui, Zihan Guan, Vishal G. Reddy, Joshua C. Eby, Gregory Madden, Rupesh Silwal, Anil Vullikanti

机构 * Department of Computer Science, University of Virginia(大学计算机科学系) University of Virginia School of Medicine(弗吉尼亚大学医学院) Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) Biocomplexity Institute and Initiative, University of Virginia(大学生物复杂性研究所) Division of Infectious Diseases & International Health, University of Virginia School of Medicine(大学感染性疾病与国际卫生分会)

专题命中 多模态RAG :knowledge retrieval(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11974 2025-07-17 cs.RO 50%

A Review of Generative AI in Aquaculture: Foundations, Applications, and Future Directions for Smart and Sustainable Farming

Waseem Akram, Muhayy Ud Din, Lyes Saad Soud, Irfan Hussain

机构 * Khalifa University Center for Autonomous Robotic Systems (KUCARS), Khalifa University, United Arab Emirates(卡里法大学自主机器人系统中心(KUCARS)、卡里法大学、阿拉伯联合酋长国)

专题命中 多模态RAG :retrieval augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02900 2025-07-11 cs.CV 50%

MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

Yunfei Xie, Ce Zhou, Lang Gao, Juncheng Wu, Xianhang Li, Hong-Yu Zhou, Sheng Liu, Lei Xing, James Zou, Cihang Xie, Yuyin Zhou

机构 * Huazhong University of Science and Technology(华中科技大学) UC Santa Cruz(加州大学圣克ruz分校) Harvard University(哈佛大学) Stanford University(斯坦福大学)

专题命中 多模态RAG :retrieval-augmented generation(abstract)

Comments The dataset is publicly available at https://yunfeixie233.github.io/MedTrinity-25M/. Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18899 2025-06-24 cs.CV 50%

FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation

Kaiyi Huang, Yukun Huang, Xintao Wang, Zinan Lin, Xuefei Ning, Pengfei Wan, Di Zhang, Yu Wang, Xihui Liu

机构 * The University of Hong Kong(香港大学) Kuaishou Technology(快手科技) Microsoft Research(微软研究院) Tsinghua University(清华大学)

专题命中 多模态RAG :RAG(abstract)

Comments Project Page: https://filmaster-ai.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17837 2025-06-24 cs.CV 50%

Time-Contrastive Pretraining for In-Context Image and Video Segmentation

Assefa Wahd, Jacob Jaremko, Abhilash Hareendranathan

机构 * Department of Radiology(放射科部门)

专题命中 多模态RAG :retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10533 2025-05-16 cs.CV cs.LG 50%

Enhancing Multi-Image Question Answering via Submodular Subset Selection

Aaryan Sharma, Shivansh Gupta, Samar Agarwal, Vishak Prasad C., Ganesh Ramakrishnan

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔学院)

专题命中 多模态RAG :retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05782 2025-03-18 cs.SD cs.CV cs.LG cs.MM eess.AS 50%

Sequential Contrastive Audio-Visual Learning

Ioannis Tsiamas, Santiago Pascual, Chunghsin Yeh, Joan Serrà

专题命中 多模态RAG :hybrid retrieval(abstract)

Comments ICASSP 2025. Version 1 contains more details

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02692 2025-03-05 cs.CE econ.GN q-fin.EC 50%

FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting

Congluo Xu, Zhaobin Liu, Ziyang Li

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21068 2025-03-03 cs.SE 50%

GUIDE: LLM-Driven GUI Generation Decomposition for Automated Prototyping

Kristian Kolthoff, Felix Kretzer, Christian Bartelt, Alexander Maedche, Simone Paolo Ponzetto

专题命中 多模态RAG :retrieval-augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17038 2025-02-25 cs.MM 50%

Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction

Jiacheng Lu, Mingyuan Xiao, Weijian Wang, Yuxin Du, Zhengze Wu, Cheng Hua

专题命中 多模态RAG :retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20725 2024-12-31 cs.CV 50%

Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling

Min Zhang, Zilin Wang, Liyan Chen, Kunhong Liu, Juncong Lin

专题命中 多模态RAG :retrieval-augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09936 2024-12-16 cs.CV 50%

CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models

Dongyu Yao, Keling Yao, Junhong Zhou, Yinghao Zhang

专题命中 多模态RAG :RAG(abstract)

Comments Disclaimer: This work is part of a course project and reflects ongoing exploration in the field of vision-language models and calorie estimation. Findings and conclusions are subject to further validation and refinement

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00304 2024-11-04 cs.CV cs.MM 50%

Unified Generative and Discriminative Training for Multi-modal Large Language Models

Wei Chow, Juncheng Li, Qifan Yu, Kaihang Pan, Hao Fei, Zhiqi Ge, Shuai Yang, Siliang Tang, Hanwang Zhang, Qianru Sun

专题命中 多模态RAG :retrieval-augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19720 2024-10-01 cs.CV 50%

FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image Classification

Kexue Fu, Xiaoyuan Luo, Linhao Qu, Shuo Wang, Ying Xiong, Ilias Maglogiannis, Longxiang Gao, Manning Wang

专题命中 多模态RAG :knowledge retrieval(abstract)

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04717 2024-05-09 cs.CV 50%

Remote Diffusion

Kunal Sunil Kasodekar

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03427 2023-10-03 eess.IV cs.CV cs.LG 50%

Merging-Diverging Hybrid Transformer Networks for Survival Prediction in Head and Neck Cancer

Mingyuan Meng, Lei Bi, Michael Fulham, Dagan Feng, Jinman Kim

专题命中 多模态RAG :RAG(abstract)

Comments Early Accepted at International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2023)

Journal ref International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), pp. 400-410, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09564 2023-08-21 cs.CV 50%

Deep Equilibrium Object Detection

Shuai Wang, Yao Teng, Limin Wang

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏