arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2759 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2759 篇

2508.09736 2025-10-10 cs.CV 83%

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

Lin Long, Yichen He, Wentao Ye, Yiyuan Pan, Yuan Lin, Hang Li, Junbo Zhao, Wei Li

机构 * Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03612 2025-10-07 cs.AI cs.CR 83%

Cross-Modal Content Optimization for Steering Web Agent Preferences

Tanqiu Jiang, Min Bai, Nikolaos Pappas, Yanjun Qi, Sandesh Swamy

机构 * Stony Brook University(石溪大学) AWS AI Labs(亚马逊人工智能实验室)

专题命中 多模态Agent :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02592 2025-10-06 cs.AI 83%

Multimodal Large Language Model Framework for Safe and Interpretable Grid-Integrated EVs

Jean Douglas Carvalho, Hugo Kenji, Ahmad Mohammad Saber, Glaucia Melo, Max Mauro Dias Santos, Deepa Kundur

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments This paper has been presented at the 2025 IEEE PES Conference on Innovative Smart Grid Technologies (ISGT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00907 2025-10-06 cs.AI 83%

Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning

Ram Ramrakhya, Matthew Chang, Xavier Puig, Ruta Desai, Zsolt Kira, Roozbeh Mottaghi

机构 * Georgia Institute of Technology(佐治亚理工学院) Meta FAIR

专题命中 多模态Agent :multimodal(title);multi-modal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05557 2025-09-09 cs.AI 83%

MV-Debate: Multi-view Agent Debate with Dynamic Reflection Gating for Multimodal Harmful Content Detection in Social Media

Rui Lu, Jinhe Bi, Yunpu Ma, Feng Xiao, Yuntao Du, Yijun Tian

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02544 2025-09-08 cs.CL 83%

Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions

Xinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang, Hai Zhao

机构 * School of Computer Science(计算机学院) Key Laboratory of Shanghai Education Commission for Intelligent Interaction and Cognitive Engineering(智能交互与认知工程重点实验室) Shanghai Jiao Tong University(上海交通大学) Shanghai Key Laboratory of Trusted Data Circulation and Governance in Web3(Web3可信数据流通与治理上海市重点实验室) GenAI, Meta(Meta GenAI)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18108 2025-08-26 cs.CL 83%

SentiMM: A Multimodal Multi-Agent Framework for Sentiment Analysis in Social Media

Xilai Xu, Zilin Zhao, Chengye Song, Zining Wang, Jinhe Qiang, Jiongrui Yan, Yuhuai Lin

机构 * College of Information and Electrical Engineering, China Agricultural University(信息与电气工程学院,中国农业大学) College of Software, Jilin University(软件学院,吉林大学) College of Communication Engineering, Jilin University(通信工程学院,吉林大学)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05580 2025-08-08 cs.CV 83%

Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis

Kunyu Feng, Yue Ma, Xinhua Zhang, Boshi Liu, Yikuang Yuluo, Yinhan Zhang, Runtao Liu, Hongyu Liu, Zhiyuan Qin, Shanhui Mo, Qifeng Chen, Zeyu Wang

机构 * HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) Tsinghua Univerisity(清华大学) Peking University(北京大学) Chongqing University(重庆大学) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)

专题命中 多模态Agent :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06418 2025-07-10 q-bio.QM cs.CV stat.AP 83%

PAST: A multimodal single-cell foundation model for histopathology and spatial transcriptomics in cancer

Changchun Yang, Haoyang Li, Yushuai Wu, Yilan Zhang, Yifeng Jiao, Yu Zhang, Rihan Huang, Yuan Cheng, Yuan Qi, Xin Guo, Xin Gao

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19835 2025-06-25 cs.CL 83%

MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Yucheng Zhou, Lingran Song, Jianbing Shen

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(abstract);分类 cs.CL

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00123 2025-06-03 cs.CV cs.RO 83%

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Gen Luo, Ganlin Yang, Ziyang Gong, Guanzhou Chen, Haonan Duan, Erfei Cui, Ronglei Tong, Zhi Hou, Tianyi Zhang, Zhe Chen, Shenglong Ye, Lewei Lu, Jingbo Wang, Wenhai Wang, Jifeng Dai, Yu Qiao, Rongrong Ji, Xizhou Zhu

机构 * Shanghai AI Laboratory(上海人工智能实验室) Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学) Xiamen University(厦门大学) SenseTime Research(商汤科技研究院) Zhejiang University(浙江大学) Nanjing University(南京大学)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22241 2025-05-29 cs.AI 83%

Agent-Centric Personalized Multiple Clustering with Multi-Modal LLMs

Ziye Chen, Yiqun Duan, Riheng Zhu, Zhenbang Sun, Mingming Gong

机构 * TikTok, Australia School of Mathematics and Statistics, University of Melbourne(墨尔本大学数学与统计学学院)

专题命中 多模态Agent :multi-modal(title,abstract);image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02937 2025-05-27 cs.CL 83%

Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent

Yangning Li, Yinghui Li, Xinyu Wang, Yong Jiang, Zhen Zhang, Xinran Zheng, Hui Wang, Hai-Tao Zheng, Philip S. Yu, Fei Huang, Jingren Zhou

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Peng Cheng Laboratory(鹏城实验室) Tongyi Lab, Alibaba Group(阿里云实验室) University College London(伦敦大学学院) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18132 2025-05-21 cs.CL 83%

MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection

Yibo Yan, Shen Wang, Jiahao Huo, Philip S. Yu, Xuming Hu, Qingsong Wen

机构 * Squirrel Ai Learning The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 多模态Agent :multimodal(title,abstract);image-text(abstract);分类 cs.CL

Comments Accepted by The 63rd Annual Meeting of the Association for Computational Linguistics (ACL Industry 2025, Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09787 2025-05-16 cs.AI 83%

A Multimodal Multi-Agent Framework for Radiology Report Generation

Ziruo Yi, Ting Xiao, Mark V. Albert

机构 * University of North Texas(北卡罗来纳州立大学)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08747 2025-04-15 cs.AI cs.IR 83%

GridMind: A Multi-Agent NLP Framework for Unified, Cross-Modal NFL Data Insights

Jordan Chipka, Chris Moyer, Clay Troyer, Tyler Fuelling, Jeremy Hochstedler

专题命中 多模态Agent :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments 16 pages, 2 figures, submitted to 2025 Sloan Sports Analytics Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04789 2025-04-08 cs.AI 83%

Multimodal Agricultural Agent Architecture (MA3): A New Paradigm for Intelligent Agricultural Decision-Making

Zhuoning Xu, Jian Xu, Mingqing Zhang, Peijie Wang, Chao Deng, Cheng-Lin Liu

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08308 2025-03-12 cs.AI 83%

Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework

Zhuo Zhi, Chen Feng, Adam Daneshmend, Mine Orlu, Andreas Demosthenous, Lu Yin, Da Li, Ziquan Liu, Miguel R. D. Rodrigues

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19902 2025-03-12 cs.AI 83%

Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy

Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Dongmei Jiang, Liqiang Nie

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Accept to CVPR 2025, Project page: https://cybertronagent.github.io/Optimus-2.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04408 2025-02-10 cs.LG cs.AI 83%

Transforming Multimodal Models into Action Models for Radiotherapy

Matteo Ferrante, Alessandra Carosi, Rolando Maria D Angelillo, Nicola Toschi

专题命中 多模态Agent :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14189 2025-01-27 cs.AI cs.LG cs.MA 83%

Distributed Multi-Agent Coordination Using Multi-Modal Foundation Models

Saaduddin Mahmud, Dorian Benhamou Goldfajn, Shlomo Zilberstein

专题命中 多模态Agent :multi-modal(title);multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10840 2024-12-17 cs.CV 83%

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning

Hai-Ming Xu, Qi Chen, Lei Wang, Lingqiao Liu

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05555 2024-12-10 cs.SE cs.AI 83%

Fragmented Layer Grouping in GUI Designs Through Graph Learning Based on Multimodal Information

Yunnong Chen, Shuhong Xiao, Jiazhi Li, Tingting Zhou, Yanfang Chang, Yankun Zhen, Lingyun Sun, Liuqing Chen

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments 28 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11530 2024-07-23 cs.CV 83%

Efficient Multimodal Learning from Data-centric Perspective

Muyang He, Yexin Liu, Boya Wu, Jianhao Yuan, Yueze Wang, Tiejun Huang, Bo Zhao

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00092 2024-07-02 cs.AI cs.ET cs.GT cs.MA 83%

Visual Reasoning and Multi-Agent Approach in Multimodal Large Language Models (MLLMs): Solving TSP and mTSP Combinatorial Challenges

Mohammed Elhenawy, Ahmad Abutahoun, Taqwa I. Alhadidi, Ahmed Jaber, Huthaifa I. Ashqar, Shadi Jaradat, Ahmed Abdelhay, Sebastien Glaser, Andry Rakotonirainy

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03568 2024-01-29 cs.AI cs.HC cs.LG 83%

Agent AI: Surveying the Horizons of Multimodal Interaction

Zane Durante, Qiuyuan Huang, Naoki Wake, Ran Gong, Jae Sung Park, Bidipta Sarkar, Rohan Taori, Yusuke Noda, Demetri Terzopoulos, Yejin Choi, Katsushi Ikeuchi, Hoi Vo, Li Fei-Fei, Jianfeng Gao

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03945 2024-01-09 cs.CL 83%

SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems

Dong Zhang, Zhaowei Li, Pengyu Wang, Xin Zhang, Yaqian Zhou, Xipeng Qiu

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(abstract);分类 cs.CL

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16217 2023-12-29 cs.CV cs.RO 83%

ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation

Xiaoqi Li, Mingxu Zhang, Yiran Geng, Haoran Geng, Yuxing Long, Yan Shen, Renrui Zhang, Jiaming Liu, Hao Dong

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08640 2023-06-29 cs.CV 83%

AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn

Difei Gao, Lei Ji, Luowei Zhou, Kevin Qinghong Lin, Joya Chen, Zihan Fan, Mike Zheng Shou

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Project page: https://showlab.github.io/assistgpt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.03378 2023-03-07 cs.LG cs.AI cs.RO 83%

PaLM-E: An Embodied Multimodal Language Model

Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, Pete Florence

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏