arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

2511.13476 2025-11-18 cs.AI 79%

Multi-Agent Multimodal Large Language Model Framework for Automated Interpretation of Fuel Efficiency Analytics in Public Transportation

Zhipeng Ma, Ali Rida Bahja, Andreas Burgdorf, André Pomp, Tobias Meisen, Bo Nørregaard Jørgensen, Zheng Grace Ma

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Journal ref Applied Sciences, 2025, 15(21), 11619

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12586 2025-11-18 cs.CL 79%

MMWOZ: Building Multimodal Agent for Task-oriented Dialogue

Pu-Hai Yang, Heyan Huang, Heng-Da Xu, Fanshu Sun, Xian-Ling Mao, Chaoxu Mu

机构 * School of Artificial Intelligence, Anhui University, Hefei, China(人工智能学院,安徽大学,合肥,中国) School of Computer Science(计算机科学学院) School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China(计算机科学与技术学院,北京理工大学,北京,中国)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19768 2025-11-18 cs.CL 79%

T^2Agent A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree Search

Xing Cui, Yueying Zou, Zekun Li, Peipei Li, Xinyuan Xu, Xuannan Liu, Huaibo Huang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

Comments accepted by AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11257 2025-11-17 cs.AI cs.CE cs.LG 79%

AIonopedia: an LLM agent orchestrating multimodal learning for ionic liquid discovery

Yuqi Yin, Yibo Fu, Siyuan Wang, Peng Sun, Hongyu Wang, Xiaohui Wang, Lei Zheng, Zhiyong Li, Zhirong Liu, Jianji Wang, Zhaoxi Sun

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09804 2025-11-14 cs.AI 79%

SlideBot: A Multi-Agent Framework for Generating Informative, Reliable, Multi-Modal Presentations

Eric Xie, Danielle Waterfield, Michael Kennedy, Aidong Zhang

专题命中 多模态Agent :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments 32 pages, 14 figures, accepted into EAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26114 2025-10-31 cs.CV 79%

OracleAgent: A Multimodal Reasoning Agent for Oracle Bone Script Research

Caoshuo Li, Zengmao Ding, Xiaobin Hu, Bang Li, Donghao Luo, Xu Peng, Taisong Jin, Yongge Liu, Shengwei Han, Jing Yang, Xiaoping He, Feng Gao, AndyPian Wu, SevenShu, Chaoyang Wang, Chengjie Wang

机构 * Xiamen University(厦门大学) Anyang Normal University(安阳师范学院) Tencent YouTu Lab(腾讯YouTu实验室) Tencent SSV(腾讯SSV)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21142 2025-10-29 cs.AI cs.LG q-bio.NC 79%

Multimodal Dreaming: A Global Workspace Approach to World Model-Based Reinforcement Learning

Léopold Maytié, Roland Bertin Johannet, Rufin VanRullen

机构 * Univ Toulouse, CNRS, CerCo, and ANITI, Artificial and Natural Intelligence Toulouse Institute(图卢兹大学、法国国家科学研究中心、CerCo以及ANITI人工智能与自然智能图卢兹研究所)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23691 2025-10-29 cs.AI 79%

Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents

Zihao Wang, Xujing Li, Yining Ye, Junjie Fang, Haoming Wang, Longxiang Liu, Shihao Liang, Junting Lu, Zhiyong Wu, Jiazhan Feng, Wanjun Zhong, Zili Li, Yu Wang, Yu Miao, Bo Zhou, Yuanfan Li, Hao Wang, Zhongkai Zhao, Faming Wu, Zhengxuan Jiang, Weihao Tan, Heyuan Yao, Shi Yan, Xiangyang Li, Yitao Liang, Yujia Qin, Guang Shi

机构 * Bytedance Seed(字节跳动种子)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20077 2025-10-29 cs.RO cs.CV cs.HC 79%

Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning

Xun Li, Rodrigo Santa Cruz, Mingze Xi, Hu Zhang, Madhawa Perera, Ziwei Wang, Ahalya Ravendran, Brandon J. Matthews, Feng Xu, Matt Adcock, Dadong Wang, Jiajun Liu

机构 * CSIRO(澳大利亚联邦科学与工业研究组织)

专题命中 多模态Agent :multi-modal(title);multimodal(abstract);分类 cs.CV

Journal ref MM '25: Proceedings of the 33rd ACM International Conference on Multimedia (2025) Pages 12492 - 12500

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19497 2025-10-23 cs.MA cs.AI 79%

Modeling realistic human behavior using generative agents in a multimodal transport system: Software architecture and Application to Toulouse

Trung-Dung Vu, Benoit Gaudou, Kamaldeep Singh Oberoi

机构 * UMR IRIT University Toulouse Capitole(IRIT大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15261 2025-10-20 cs.AI 79%

AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory

Jitesh Jain, Shubham Maheshwari, Ning Yu, Wen-mei Hwu, Humphrey Shi

机构 * Adobe(Adobe公司) Netflix Eyeline Studios(Netflix Eyeline工作室) UIUC(伊利诺伊大学香槟分校)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments LAW 2025 Workshop at NeurIPS 2025. Work done from late 2023 to early 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19905 2025-10-17 cs.AI 79%

EMAC+: Embodied Multimodal Agent for Collaborative Planning with VLM+LLM

Shuang Ao, Flora D. Salim, Simon Khan

机构 * University of New South Wales(新南威尔士大学) Air Force Research Laboratory(空军研究实验室)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09762 2025-10-14 cs.LG cs.AI 79%

PatentVision: A multimodal method for drafting patent applications

Ruo Yang, Sai Krishna Reddy Mudhiganti, Manali Sharma

机构 * Samsung Semiconductor, Inc.(三星半导体公司)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06223 2025-10-10 cs.HC cs.AI 79%

A Multimodal GUI Architecture for Interfacing with LLM-Based Conversational Assistants

Hans G. W. van Dam

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 24 pages, 19 figures, code available at https://github.com/hansvdam/langbar

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06664 2025-10-09 cs.CL 79%

ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory

Yunzhong Xiao, Yangmin Li, Hewei Wang, Yunlong Tang, Zora Zhiruo Wang

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Rochester(罗切斯特大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04560 2025-10-07 cs.AI 79%

ContextNav: Towards Agentic Multimodal In-Context Learning

Honghao Fu, Yuan Ouyang, Kai-Wei Chang, Yiwei Wang, Zi Huang, Yujun Cai

机构 * The University of Queensland(昆士兰大学) Nanjing University(南京大学) University of California, Los Angeles(加州大学洛杉矶分校) University of California, Merced(加州大学默塞德分校)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02677 2025-10-06 cs.AI cs.LG 79%

ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks

Zhaorun Chen, Xun Liu, Mintong Kang, Jiawei Zhang, Minzhou Pan, Shuang Yang, Bo Li

机构 * University of Chicago(芝加哥大学) University of Illinois(伊利诺伊大学) Virtue AI Meta

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 60 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00496 2025-10-06 cs.CL 79%

Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations

Pengzhou Cheng, Lingzhong Dong, Zeng Wu, Zongru Wu, Xiangru Tang, Chengwei Qin, Zhuosheng Zhang, Gongshen Liu

机构 * Shanghai Jiao Tong University(上海交通大学) Yale University(耶鲁大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

Comments 23 pages, 10 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01717 2025-10-03 cs.LG cs.AI 79%

Latency-aware Multimodal Federated Learning over UAV Networks

Shaba Shaon, Dinh C. Nguyen

机构 * ECE Department, University of Alabama in Huntsville(阿拉巴马大学亨茨维尔分校电子与计算机工程系)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments Accepted at IEEE Transactions on Network Science and Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00161 2025-10-02 cs.CL 79%

TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Ken Fukuda, Teruko Mitamura

机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所) National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

Comments 21 pages. Code: https://github.com/kimihiroh/tama

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26161 2025-10-01 cs.AI cs.SE 79%

90% Faster, 100% Code-Free: MLLM-Driven Zero-Code 3D Game Development

Runxin Yang, Yuxuan Wan, Shuqing Li, Michael R. Lyu

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态Agent :MLLM(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24855 2025-09-30 cs.AI 79%

PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System

Fangchen Yu, Junchi Yao, Ziyi Wang, Haiyuan Wan, Youling Huang, Bo Zhang, Shuyue Hu, Dongzhan Zhou, Ning Ding, Ganqu Cui, Lei Bai, Wanli Ouyang, Peng Ye

机构 * Shanghai AI Laboratory(上海人工智能实验室) CUHK-Shenzhen(香港中文大学(深圳)) CUHK(香港中文大学) UESTC(电子科技大学) Tsinghua University(清华大学) DUT(大连理工大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24314 2025-09-30 cs.AI 79%

MedMMV: A Controllable Multimodal Multi-Agent Framework for Reliable and Verifiable Clinical Reasoning

Hongjun Liu, Yinghao Zhu, Yuhui Wang, Yitao Long, Zeyu Lai, Lequan Yu, Chen Zhao

机构 * New York University(纽约大学) NYU Shanghai(纽约大学上海分校) The University of Hong Kong(香港大学) Zhejiang University(浙江大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 25 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12839 2025-09-30 cs.AI 79%

From An LLM Swarm To A PDDL-Empowered HIVE: Planning Self-Executed Instructions In A Multi-Modal Jungle

Kaustubh Vyas, Damien Graux, Yijun Yang, Sébastien Montella, Chenxin Diao, Wendi Zhou, Pavlos Vougiouklis, Ruofei Lai, Yang Ren, Keshuang Li, Jeff Z. Pan

机构 * Huawei Technologies Ltd., UK(华为技术有限公司,英国) University of Edinburgh, UK(爱丁堡大学,英国)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

Comments Published as a conference paper at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08524 2025-09-29 cs.HC cs.AI 79%

StreetReaderAI: Making Street View Accessible Using Context-Aware Multimodal AI

Jon E. Froehlich, Alexander Fiannaca, Nimer Jaber, Victor Tsaran, Shaun Kane

机构 * Google Research(谷歌研究) Google DeepMind(谷歌DeepMind) Google(谷歌)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments Accepted to UIST'25; v2. Fixed a missing word in the PDF; v3. Fixed a typo in an author's name; v4. Changed system name and title

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10547 2025-09-26 cs.RO cs.AI 79%

Real-Time Out-of-Distribution Failure Prevention via Multi-Modal Reasoning

Milan Ganai, Rohan Sinha, Christopher Agia, Daniel Morton, Luigi Di Lillo, Marco Pavone

机构 * Stanford University(斯坦福大学) Swiss Re(瑞士再保险集团) NVIDIA Research(NVIDIA研究)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

Comments Conference on Robot Learning (CoRL) 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19087 2025-09-24 cs.CV 79%

Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications

Ganesh Mallya, Yotam Gigi, Dahun Kim, Maxim Neumann, Genady Beryozkin, Tomer Shekel, Anelia Angelova

机构 * Google DeepMind(谷歌DeepMind) Google Research(谷歌研究)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16051 2025-09-22 cs.AI 79%

MMAPG: A Training-Free Framework for Multimodal Multi-hop Question Answering via Adaptive Planning Graphs

Yiheng Hu, Xiaoyang Wang, Qing Liu, Xiwei Xu, Qian Fu, Wenjie Zhang, Liming Zhu

机构 * University of New South Wales(新南威尔士大学) CSIRO Data61(CSIRO数据61)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20474 2025-09-22 q-fin.TR cs.CL cs.LG 79%

MountainLion: A Multi-Modal LLM-Based Agent System for Interpretable and Adaptive Financial Trading

Siyi Wu, Junqiao Wang, Zhaoyang Guan, Leyi Zhao, Xinyuan Song, Xinyu Ying, Dexu Yu, Jinhao Wang, Hanlin Zhang, Michele Pak, Yangfan He, Yi Xin, Jianhui Wang, Tianyu Shi

机构 * The University of Texas at Arlington(德克萨斯大学阿灵顿分校) Sichuan University(四川大学) Northwestern University(西北大学) Indiana University(印第安纳大学) Emory University(埃默里大学) Nankai University(南开大学) MountainLion Research(MountainLion研究机构) Xi’an University of Electronic Science and Technology(西安电子科技大学) Kyoto University(京都大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Nanjing university(南京大学) Tsinghua University(清华大学) University of Toronto(多伦多大学)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03603 2025-09-22 cs.AI cs.LG 79%

Towards deployment-centric multimodal AI beyond vision and language

Xianyuan Liu, Jiayang Zhang, Shuo Zhou, Thijs L. van der Plas, Avish Vijayaraghavan, Anastasiia Grishina, Mengdie Zhuang, Daniel Schofield, Christopher Tomlinson, Yuhan Wang, Ruizhe Li, Louisa van Zeeland, Sina Tabakhi, Cyndie Demeocq, Xiang Li, Arunav Das, Orlando Timmerman, Thomas Baldwin-McDonald, Jinge Wu, Peizhen Bai, Zahraa Al Sahili, Omnia Alwazzan, Thao N. Do, Mohammod N. I. Suvon, Angeline Wang, Lucia Cipolina-Kun, Luigi A. Moretti, Lucas Farndale, Nitisha Jain, Natalia Efremova, Yan Ge, Marta Varela, Hak-Keung Lam, Oya Celiktutan, Ben R. Evans, Alejandro Coca-Castro, Honghan Wu, Zahraa S. Abdallah, Chen Chen, Valentin Danchev, Nataliya Tkachenko, Lei Lu, Tingting Zhu, Gregory G. Slabaugh, Roger K. Moore, William K. Cheung, Peter H. Charlton, Haiping Lu

机构 * Centre for Machine Intelligence, University of Sheffield, Sheffield, UK(机器智能中心,谢菲尔德大学) School of Computer Science, University of Sheffield, Sheffield, UK(计算机科学学院,谢菲尔德大学) The Alan Turing Institute, London, UK(艾伦·图灵研究所,伦敦) Department of Metabolism, Digestion and Reproduction, Imperial College London, London, UK(代谢、消化与生殖部门,伦敦帝国学院) Department of Applied AI, Simula Research Laboratory, Oslo, Norway(应用人工智能部门,Simula研究实验室,奥斯陆) Information School, University of Sheffield, Sheffield, UK(信息学院,谢菲尔德大学) NHS England, Leeds, UK(英国英格兰国家医疗服务体系,利兹) Institute of Health Informatics, University College London, London, UK(健康信息研究所,伦敦大学学院) Department of Engineering, King’s College London, London, UK(工程部门,伦敦国王学院) Department of Computing Science, University of Aberdeen, Aberdeen, UK(计算科学部门,阿伯丁大学) School of Informatics, University of Edinburgh, Edinburgh, UK(信息学院,爱丁堡大学) School of Engineering Mathematics and Technology, University of Bristol, Bristol, UK(工程数学与技术学院,布里斯托尔大学) Department of Informatics, King’s College London, London, UK(信息部门,伦敦国王学院) Department of Earth Sciences, University of Cambridge, Cambridge, UK(地球科学部门,剑桥大学) Department of Computer Science, University of Manchester, Manchester, UK(计算机科学部门,曼彻斯特大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏