arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

2502.16865 2025-02-25 cs.IR 82%

Multimodal Search in Chemical Documents and Reactions

Ayush Kumar Shah, Abhisek Dey, Leo Luo, Bryan Amador, Patrick Philippy, Ming Zhong, Siru Ouyang, David Mark Friday, David Bianchi, Nick Jackson, Richard Zanibbi, Jiawei Han

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract)

Comments 4 pages, 2 figures, SIGIR 2025 Demonstration Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14528 2025-02-21 math.OC 82%

Dynamic Preference-based Multi-modal Trip Planning of Public Transport and Shared Mobility

Yimeng Zhang, Oded Cats, Shadi Sharif Azadeh

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00252 2025-02-18 cs.AI cs.CL cs.CV cs.MA 82%

Towards Rationality in Language and Multimodal Agents: A Survey

Bowen Jiang, Yangxinyu Xie, Xiaomeng Wang, Yuan Yuan, Zhuoqun Hao, Xinyi Bai, Weijie J. Su, Camillo J. Taylor, Tanwi Mallick

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments This paper has been accepted to the NAACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14394 2025-02-13 cs.AI cs.CL cs.CV 82%

A Multimodal Automated Interpretability Agent

Tamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram, Evan Hernandez, Jacob Andreas, Antonio Torralba

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 25 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12574 2025-01-24 cs.AI cs.CL cs.CV cs.LG 82%

MuMA-ToM: Multi-modal Multi-Agent Theory of Mind

Haojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin, Leyla Isik, Yen-Ling Kuo, Tianmin Shu

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments AAAI-25 (Oral). Project website: https://scai.cs.jhu.edu/projects/MuMA-ToM/ Code: https://github.com/SCAI-JHU/MuMA-ToM

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11051 2025-01-22 cs.CV cs.AI cs.CL cs.RO 82%

FLAME: Learning to Navigate with Multimodal LLM in Urban Environments

Yunzhe Xu, Yiyuan Pan, Zhe Liu, Hesheng Wang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to AAAI 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11974 2024-12-18 cs.RO cs.AI cs.CL cs.CV 82%

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Qi Sun, Pengfei Hong, Tej Deep Pala, Vernon Toh, U-Xuan Tan, Deepanway Ghosal, Soujanya Poria

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments https://github.com/declare-lab/Emma-X, https://huggingface.co/declare-lab/Emma-X

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08442 2024-12-12 cs.LG 82%

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

Andrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, Devon Hjelm, Zhe Gan, Zsolt Kira, Alexander Toshev

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07904 2024-12-10 cs.LG 82%

Grounding Multimodal Large Language Models in Actions

Andrew Szot, Bogdan Mazoure, Harsh Agrawal, Devon Hjelm, Zsolt Kira, Alexander Toshev

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10603 2024-11-19 cs.RO cs.SY eess.SY 82%

A Novel MLLM-based Approach for Autonomous Driving in Different Weather Conditions

Sonda Fourati, Wael Jaafar, Noura Baccar

专题命中 多模态Agent :MLLM(title,abstract);multi-modal(abstract)

Comments 9 pages, 6 figures; Submitted to IEEE Transactions on Intelligent Transportation Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21480 2024-10-30 cs.LG cs.AI cs.CL cs.CV 82%

AiSciVision: A Framework for Specializing Large Multimodal Models in Scientific Image Classification

Brendan Hogan, Anmol Kabra, Felipe Siqueira Pacheco, Laura Greenstreet, Joshua Fan, Aaron Ferber, Marta Ummus, Alecsander Brito, Olivia Graham, Lillian Aoki, Drew Harvell, Alex Flecker, Carla Gomes

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14277 2024-09-24 cs.AI cs.CL cs.CV cs.RO 82%

Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models

Yew Ken Chia, Qi Sun, Lidong Bing, Soujanya Poria

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06327 2024-08-13 cs.AI cs.CL cs.CV 82%

VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Yifan Xu, Xixuan Song, Shudan Zhang, Hanyu Lai, Xinyi Liu, Hanlin Zhao, Jiadai Sun, Xinyue Yang, Yu Yang, Zehan Qi, Shuntian Yao, Xueqiao Sun, Siyi Cheng, Qinkai Zheng, Hao Yu, Hanchen Zhang, Wenyi Hong, Ming Ding, Lihang Pan, Xiaotao Gu, Aohan Zeng, Zhengxiao Du, Chan Hee Song, Yu Su, Yuxiao Dong, Jie Tang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14972 2024-08-07 cs.AI cs.CL cs.MA cs.MM 82%

A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning

Changmeng Zheng, Dayong Liang, Wengyu Zhang, Xiao-Yong Wei, Tat-Seng Chua, Qing Li

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted by ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02121 2024-08-06 physics.flu-dyn physics.app-ph physics.comp-ph 82%

Non-invasive imaging assisted CFD simulation of 4D multi-modal fluid flow using In-situ adaptor

Vaishali Sharma, Arpit Kumar, Snehlata Shakya, Mayank Goswami

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(abstract)

Comments 11 Pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18035 2024-07-26 cs.CV cs.AI cs.CL 82%

RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models

Haoyu Chen, Wenbo Li, Jinjin Gu, Jingjing Ren, Sixiang Chen, Tian Ye, Renjing Pei, Kaiwen Zhou, Fenglong Song, Lei Zhu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01587 2024-06-05 cs.RO 82%

PlanAgent: A Multi-modal Large Language Agent for Closed-loop Vehicle Motion Planning

Yupeng Zheng, Zebin Xing, Qichao Zhang, Bu Jin, Pengfei Li, Yuhang Zheng, Zhongpu Xia, Kun Zhan, Xianpeng Lang, Yaran Chen, Dongbin Zhao

专题命中 多模态Agent :multi-modal(title,abstract);MLLM(abstract)

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18358 2024-05-29 cs.CL cs.AI cs.CV cs.LG 82%

MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning

Somnath Kumar, Yash Gadhia, Tanuja Ganu, Akshay Nambi

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16829 2024-05-27 cs.CV cs.AI cs.CL 82%

Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials

Ye Fang, Zeyi Sun, Tong Wu, Jiaqi Wang, Ziwei Liu, Gordon Wetzstein, Dahua Lin

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://sunzey.github.io/Make-it-Real/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18137 2024-05-27 cs.RO cs.AI cs.CL cs.CV cs.LG 82%

DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning

Jianxiong Li, Jinliang Zheng, Yinan Zheng, Liyuan Mao, Xiao Hu, Sijie Cheng, Haoyi Niu, Jihao Liu, Yu Liu, Jingjing Liu, Ya-Qin Zhang, Xianyuan Zhan

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11640 2024-05-21 cs.AI cs.CL cs.CV 82%

Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning

Zishan Gu, Fenglin Liu, Changchang Yin, Ping Zhang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04950 2024-05-09 cs.CV cs.AI cs.CL 82%

VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context

Yunxin Li, Baotian Hu, Haoyuan Shi, Wei Wang, Longyue Wang, Min Zhang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 17 pages; Accepted by ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08268 2023-10-12 cs.RO cs.AI cs.CL cs.LG cs.SD eess.AS 82%

Chat with the Environment: Interactive Multimodal Perception Using Large Language Models

Xufeng Zhao, Mengdi Li, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments IROS2023, Detroit. See the project website at https://matcha-agent.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.07363 2020-11-17 cs.IR 82%

RecTen: A Recursive Hierarchical Low Rank Tensor Factorization Method to Discover Hierarchical Patterns in Multi-modal Data

Risul Islam, Md Omar Faruk Rokon, Evangelos E. Papalexakis, Michalis Faloutsos

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(abstract)

Comments 9 pages, 9 figures, 1 table, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14593 2026-07-20 cs.HC cs.AI cs.CL 版本更新 82%

Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction

记忆驱动的自我表露与关系转折点:人机交互的纵向多模态研究

Ryuichi Sumida, Mao Saeki, Masaki Eguchi, Sadahiro Yoshikawa, Koji Inoue, Tatsuya Kawahara, Yoichi Matsuyama

机构 * Graduate School of Informatics, Kyoto University(京都大学信息学研究科) Equmenopolis, Inc.(Equmenopolis公司) Waseda University(早稻田大学) School of Informatics, Kyoto University(京都大学信息学系)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 该研究通过纵向多模态研究,探讨对话式人工智能系统中交互如何发展为关系。核心方法是让参与者对五个关系构建要素评分,发现对话质量影响当下愉悦感,感知记忆受关系制约,关系有崩溃和激增等转折点,揭示了人机关系建立的方式。

Comments 15 pages, 3 figures. Accepted to ICMI 2026 (International Conference on Multimodal Interaction), October 5-9, 2026, Napoli, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02629 2025-08-07 cs.RO cs.AI cs.CL 82%

HyCodePolicy: Hybrid Language Controllers for Multimodal Monitoring and Decision in Embodied Agents

Yibin Liu, Zhixuan Liang, Zanxin Chen, Tianxing Chen, Mengkang Hu, Wanxi Dong, Congsheng Xu, Zhaoming Han, Yusen Qin, Yao Mu

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI;multi-modal(comments)

Comments Accepted to ICCV 2025 Workshop on Multi-Modal Reasoning for Agentic Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16170 2023-12-27 cs.CV cs.AI cs.RO 82%

EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

Tai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu, Ruiyuan Lyu, Peisen Li, Xiao Chen, Wenwei Zhang, Kai Chen, Tianfan Xue, Xihui Liu, Cewu Lu, Dahua Lin, Jiangmiao Pang

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments A multi-modal, ego-centric 3D perception dataset and benchmark for holistic 3D scene understanding. Project page: http://tai-wang.github.io/embodiedscan

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11844 2026-08-05 cs.CV 版本更新 81%

Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

超越单摄像头:体育视频理解中的智能多视角推理

Kerui Chen, Jinglu Wang, Xiaoyi Zhang, Yan Lu

机构 * Zhejiang University(浙江大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态Agent :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV

AI总结 针对体育视频多视角理解缺乏评估基准及MLLMs难以利用多视角信息的问题,引入SportMV - Bench基准,分析瓶颈所在,并提出SportMV - Agent框架,通过迭代循环实现主动视角选择等,相比最强MLLM基线有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16300 2026-07-23 cs.AI 版本更新 81%

Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection

循环代码取证:面向图像伪造检测的代理工具使用

Fanrui Zhang, Qiang Zhang, Sizhuo Zhou, Jianwen Sun, Chuanhao Li, Jiaxin Ai, Yukang Feng, Yujie Zhang, Wenjie Li, Zizhen Li, Yifan Chang, Jiawei Liu, Kaipeng Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态Agent :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 本文提出ForenAgent框架,通过多轮交互使MLLM自主生成并迭代优化Python工具,提升图像伪造检测的灵活性和可解释性,构建FABench数据集进行系统训练和评估。

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09290 2026-06-09 cs.CV 新提交 81%

Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning

Visual Para-Thinker++:用于视觉推理的单策略多智能体框架

Haoran Xu, Hongyu Wang, Yifei Gao, Jiaze Li, Zizhao Tong, Xiaofeng Zhang, Xiaosong Yuan

机构 * Zhejiang University(浙江大学) Hunan University(湖南大学) Tianjin University(天津大学) University of Chinese Academy of Sciences(中国科学院大学) Shanghai Jiao Tong University(上海交通大学) Jilin University(吉林大学)

专题命中 多模态Agent :MLLM(summary_cn,abstract);分类 cs.CV

AI总结 提出Visual Para-Thinker++框架,通过共享MLLM策略实例化为多个角色智能体并行推理,结合多智能体能力注入和角色解耦优化,有效缓解视觉推理中的早期感知承诺和幻觉问题。

详情

展开后加载摘要…

URL PDF HTML 收藏