arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

2401.13919 2024-06-10 cs.CL cs.AI 81%

WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, Dong Yu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ACL 2024 (main). Code and data is released at https://github.com/MinorJerry/WebVoyager

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11436 2024-06-10 cs.CL cs.AI cs.HC 81%

You Only Look at Screens: Multimodal Chain-of-Action Agents

Zhuosheng Zhang, Aston Zhang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Findings of ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13034 2024-06-07 cs.CL cs.AI cs.HC 81%

Autonomous Workflow for Multimodal Fine-Grained Training Assistants Towards Mixed Reality

Jiahuan Pei, Irene Viola, Haochen Huang, Junxiao Wang, Moonisa Ahsan, Fanghua Ye, Jiang Yiming, Yao Sai, Di Wang, Zhumin Chen, Pengjie Ren, Pablo Cesar

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19267 2024-05-24 cs.CL cs.AI 81%

MineLand: Simulating Large-Scale Multi-Agent Interactions with Limited Multimodal Senses and Physical Needs

Xianhao Yu, Jiaqi Fu, Renjia Deng, Wenjuan Han

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Project website: https://github.com/cocacola-lab/MineLand

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12222 2024-05-22 eess.IV cs.AI cs.CV 81%

Influence based explainability of brain tumors segmentation in multimodal Magnetic Resonance Imaging

Tommaso Torda, Andrea Ciardiello, Simona Gargiulo, Greta Grillo, Simone Scardapane, Cecilia Voena, Stefano Giagu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03627 2024-04-29 cs.CL cs.AI 81%

Multimodal Large Language Models to Support Real-World Fact-Checking

Jiahui Geng, Yova Kementchedjhieva, Preslav Nakov, Iryna Gurevych

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11330 2024-04-24 cs.CL cs.AI cs.HC cs.LG 81%

Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback

Dong Won Lee, Hae Won Park, Yoon Kim, Cynthia Breazeal, Louis-Philippe Morency

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06682 2024-03-12 cs.CL cs.CV cs.CY 81%

Restoring Ancient Ideograph: A Multimodal Multitask Neural Network Approach

Siyu Duan, Jun Wang, Qi Su

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accept by Lrec-Coling 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.11029 2024-03-04 cs.CL cs.AI 81%

META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, Kai Yu

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02330 2024-02-23 cs.CV cs.CL 81%

LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Yichen Zhu, Minjie Zhu, Ning Liu, Zhicai Ou, Xiaofeng Mou, Jian Tang

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments The datasets were incomplete as they did not include all the necessary copyrights

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03610 2024-02-07 cs.LG cs.AI cs.CL 81%

RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents

Tomoyuki Kagaya, Thong Jing Yuan, Yuxuan Lou, Jayashree Karlekar, Sugiri Pranata, Akira Kinose, Koki Oguri, Felix Wick, Yang You

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14859 2023-12-22 cs.CV cs.CL 81%

3M-TRANSFORMER: A Multi-Stage Multi-Stream Multimodal Transformer for Embodied Turn-Taking Prediction

Mehdi Fatan, Emanuele Mincato, Dimitra Pintzou, Mariella Dimiccoli

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07886 2023-12-14 cs.AI cs.CL 81%

Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI

Kai Huang, Boyuan Yang, Wei Gao

专题命中 多模态Agent :multimodal(title);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07562 2023-11-14 cs.CV cs.AI 81%

GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

An Yan, Zhengyuan Yang, Wanrong Zhu, Kevin Lin, Linjie Li, Jianfeng Wang, Jianwei Yang, Yiwu Zhong, Julian McAuley, Jianfeng Gao, Zicheng Liu, Lijuan Wang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04067 2023-11-08 cs.LG cs.AI cs.CV 81%

Multitask Multimodal Prompted Training for Interactive Embodied Task Completion

Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10790 2023-10-26 cs.LG cs.AI cs.CV cs.RO 81%

Guide Your Agent with Adaptive Multimodal Rewards

Changyeon Kim, Younggyo Seo, Hao Liu, Lisa Lee, Jinwoo Shin, Honglak Lee, Kimin Lee

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to NeurIPS 2023. Project webpage: https://sites.google.com/view/2023arp

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13850 2023-07-27 cs.LG cs.AI cs.CV cs.RO 81%

MAEA: Multimodal Attribution for Embodied AI

Vidhi Jain, Jayant Sravan Tamarapalli, Sahiti Yerramilli, Yonatan Bisk

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06485 2023-05-12 cs.RO cs.AI cs.CL cs.HC 81%

Multimodal Contextualized Plan Prediction for Embodied Task Completion

Mert İnan, Aishwarya Padmakumar, Spandana Gella, Patrick Lange, Dilek Hakkani-Tur

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments NILLI at EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.09151 2020-08-24 cs.CL cs.MM 81%

Multi-modal Cooking Workflow Construction for Food Recipes

Liangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu, Chong-Wah Ngo, Min-Yen Kan, Yu-Gang Jiang, Tat-Seng Chua

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CL、cs.MM

Comments This manuscript has been accepted at ACM MM 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.01385 2019-02-05 cs.LG cs.AI cs.CL cs.RO stat.ML 81%

Embodied Multimodal Multitask Learning

Devendra Singh Chaplot, Lisa Lee, Ruslan Salakhutdinov, Devi Parikh, Dhruv Batra

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments See https://devendrachaplot.github.io/projects/EMML for demo videos

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11954 2018-11-22 cs.CL cs.AI 81%

A Knowledge-Grounded Multimodal Search-Based Conversational Agent

Shubham Agarwal, Ondrej Dusek, Ioannis Konstas, Verena Rieser

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

Journal ref Proceedings of the 2018 EMNLP Workshop SCAI: The 2nd International Workshop on Search-Oriented Conversational AI, pages 59-66, Brussels, Belgium, October 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16774 2025-10-21 cs.LG cs.AI 80%

Learning to play: A Multimodal Agent for 3D Game-Play

Yuguang Yue, Irakli Salia, Samuel Hunt, Christopher Green, Wenzhe Shi, Jonathan J Hunt

机构 * Player2

专题命中 多模态Agent :multimodal(title);multi-modal(abstract,comments);分类 cs.AI

Comments International Conference on Computer Vision Workshop on Multi-Modal Reasoning for Agentic Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.07025 2020-09-16 cs.CV 80%

FairCVtest Demo: Understanding Bias in Multimodal Learning with a Testbed in Fair Automatic Recruitment

Alejandro Peña, Ignacio Serna, Aythami Morales, Julian Fierrez

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Comments ACM Intl. Conf. on Multimodal Interaction (ICMI). arXiv admin note: substantial text overlap with arXiv:2004.07173

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19436 2026-08-12 cs.CV cs.AI cs.LG cs.MM 版本更新 80%

VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection

VDC-Agent:当视频详细描述器通过代理自我反思而自我进化

Qiang Wang, Xinyuan Gao, Yuhang He, Jizhou Han, Jiangyang Li, SongLin Dong, Zhiheng Ma, Yihong Gong

机构 * Xi’an Jiaotong University(西安交通大学) Kuaishou Technology(快手科技) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 VDC-Agent通过自我反思机制实现视频详细描述的自我进化,利用自动生成的(描述,评分)对提升描述准确性与评分表现。

Comments Accepted to ECCV 2026. Project Page: https://vdcagent.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26775 2026-08-04 cs.LG cs.AI cs.CL cs.CV 80%

Learning to Select Visual In-Context Demonstrations

学习选择视觉上下文示例

Eugene Lee, Yu-Chi Lin, Jiajie Diao

机构 * University of Cincinnati(辛辛那提大学) University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出LSD方法,通过强化学习构建最优演示集,提升多模态大语言模型在视觉回归任务中的表现,揭示了学习选择在视觉上下文学习中的必要性。

Comments 21 pages, 12 figure, accepted to Computer Vision and Pattern Recognition Conference (CVPR) 2026 Findings Track

Journal ref In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9455-9465) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27926 2026-06-29 cs.AI cs.CL cs.CV 新提交 80%

Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing

可验证几何问题求解:求解器驱动的自动形式化与定理提出

Can Li, Ting Zhang, Junbo Zhao, Hua Huang

机构 * Beijing Normal University(北京师范大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出SD-GPS框架,通过求解器驱动的自动形式化和可验证定理提出,解决几何问题求解中神经符号方法的瓶颈,在Geometry3K和PGPS9K上超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16491 2026-06-16 cs.RO 新提交 80%

HATS: A Human-Agent Teleoperation System for Multi-Arm Data Collection

HATS:用于多臂数据收集的人-智能体遥操作系统

Zesen Lin, Jian-Jian Jiang, Haoming Cen, Xiao-Ming Wu, Dandan Zhang, Wei-Shi Zheng

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Nanyang Technological University(南洋理工大学) Imperial College London(帝国理工学院)

专题命中 多模态Agent :MLLM(summary_cn,abstract)

AI总结 提出HATS系统,由单操作员借助MLLM智能体控制两主臂和两辅助臂,实现高效多臂数据收集,性能媲美双人专家团队。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13807 2026-08-17 physics.optics cs.CV physics.med-ph 新提交 79%

Label-Free Deep-Tissue Peripheral Nerve Detection with a Handheld Multimodal OCT Probe and NerveDetNet

使用手持式多模态OCT探头和NerveDetNet进行无标记深层组织周围神经检测

Yihan Wang, Ruilin You, Shaobai Li, Jiabin Chen, Bofan Song, Anh D. Le, Rongguang Liang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 该研究开发了结合手持式多模态OCT探头与NerveDetNet的无标记框架,可在不切开组织的情况下检测皮下周围神经并解析深度,其性能优于多种基线方法,具备术中应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10756 2026-08-12 cs.RO cs.CV 新提交 79%

Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting

基于语义三维高斯溅射的开放词汇移动操作具身多模态定位

Huosen Ou, Dongni Song, Yuncong Wang, Tao Zhou, Yiding Ji

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Midea Group(美的集团) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 本文针对家庭场景开放词汇移动操作的目标定位问题,提出整合Semantic-3DGS的具身多模态框架,经50次真实机器人试验,在长 horizon、杂乱场景等任务中优于PointVLA等基线方法,提升了操作鲁棒性。

Comments 9 pages, 11 figures. Accepted to ACM Multimedia 2026 (MM '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07126 2026-08-12 cs.RO cs.AI cs.LG 版本更新 79%

RankFormer: A Propose-then-Select Transformer for Multi-Agent Multimodal Trajectory Prediction

自发现意图感知变换器用于多模态车辆轨迹预测

Diyi Liu, Zihan Niu, Tu Xu, Xingchen Zhang, Lishan Sun

专题命中 多模态Agent :multimodal(title);multi-modal(abstract);分类 cs.AI

AI总结 本文提出一种基于Transformer的多模态车辆轨迹预测模型,通过分离空间模块和轨迹生成模块提升预测性能,并通过预测残差偏移学习有序轨迹。

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏