arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2777 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2777 篇

2510.03829 2025-10-07 cs.NI cs.AI 57%

A4FN: an Agentic AI Architecture for Autonomous Flying Networks

André Coelho, Pedro Ribeiro, Helder Fontes, Rui Campos

机构 * FCT – Fundação para a Ciência e a Tecnologia, I.P.(葡萄牙科学与技术基金会)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments This paper has been accepted for presentation in the Auto ML for Zero-Touch Network Management Workshop (WS04-01) at the IEEE International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00413 2025-10-07 cs.CV 57%

PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents

Zikang Liu, Junyi Li, Wayne Xin Zhao, Dawei Gao, Yaliang Li, Ji-rong Wen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学全球化人工智能学院) Department of Data Science, City University of Hong Kong(香港城市大学数据科学系) Alibaba Group(阿里巴巴集团)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07778 2025-10-07 cs.CV 57%

A Neurosymbolic Agent System for Compositional Visual Reasoning

Yichang Xu, Gaowen Liu, Ramana Rao Kompella, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Ling Liu

机构 * Georgia Institute of Technology(佐治亚理工学院) Cisco Systems(思科系统)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20140 2025-10-07 cs.AI 57%

MAD-Sherlock: Multi-Agent Debate for Visual Misinformation Detection

Kumud Lakara, Georgia Channing, Christian Rupprecht, Juil Sock, Philip Torr, John Collomosse, Christian Schroeder de Witt

机构 * University of Oxford, Oxford, UK(牛津大学) BBC AI Research, London, UK(BBC人工智能研究) University of Surrey, Guildford, UK(萨里大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23263 2025-10-06 cs.AI 57%

GUI-PRA: Process Reward Agent for GUI Tasks

Tao Xiong, Xavier Hu, Yurun Chen, Yuhang Liu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang

机构 * Zhejiang University(浙江大学) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25297 2025-10-02 cs.SE cs.AI 57%

Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development

Yuxuan Wan, Tingshuo Liang, Jiakai Xu, Jingyu Xiao, Yintong Huo, Michael R. Lyu

机构 * The Chinese University of Hong Kong(香港中文大学) Columbia University in the City of New York(哥伦比亚大学) Singapore Management University(新加坡管理学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00510 2025-10-02 cs.CL 57%

JoyAgent-JDGenie: Technical Report on the GAIA

Jiarun Liu, Shiyue Xu, Shangkun Liu, Yang Li, Wen Liu, Min Liu, Xiaoqing Zhou, Hanmin Wang, Shilin Jia, zhen Wang, Shaohua Tian, Hanhao Li, Junbo Zhang, Yongli Yu, Peng Cao, Haofen Wang

机构 * JoyAgent-JDGenie

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00078 2025-10-02 cs.LG cs.AI cs.DC 57%

Adaptive and Resource-efficient Agentic AI Systems for Mobile and Embedded Devices: A Survey

Sicong Liu, Weiye Wu, Xiangrui Xu, Teng Li, Bowen Pang, Bin Guo, Zhiwen Yu

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25885 2025-10-01 cs.AI 57%

SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents

Ruolin Chen, Yinqian Sun, Jihang Wang, Mingyang Lv, Qian Zhang, Yi Zeng

机构 * Brain-inspired Cognitive AI Lab, Institute of Automation, Chinese Academy of Sciences(脑启发认知AI实验室,自动化研究所,中国科学院) Beijing Key Laboratory of Safe AI and Superalignment(北京安全AI与超对齐关键实验室) Beijing Institute of AI Safety and Governance(北京AI安全与治理研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Long-term AI(长期AI)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11790 2025-09-30 cs.AI 57%

Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs

Nasim Borazjanizadeh, Roei Herzig, Eduard Oks, Trevor Darrell, Rogerio Feris, Leonid Karlinsky

机构 * Xero Inc.(Xero公司) Berkeley AI Research, UC Berkeley(伯克利人工智能研究实验室,伯克利大学) MIT–IBM Watson AI Lab(麻省理工–IBM沃森人工智能实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07978 2025-09-30 cs.AI quant-ph 57%

Agents for self-driving laboratories applied to quantum computing

Shuxiang Cao, Zijian Zhang, Mohammed Alghadeer, Simone D Fasciati, Michele Piscitelli, Mustafa Bakr, Peter Leek, Alán Aspuru-Guzik

机构 * 1 Clarendon Laboratory, Department of Physics, University of Oxford, Oxford, OX1 3PU, UK 2 Department of Computer Science, University of Toronto, Toronto, ON M5S 2E4, Canada 3 Vector Institute for Artificial Intelligence, Toronto, ON, M5G 1M1, Canada 4 Department of Chemistry, University of Toronto, Toronto, ON M5S 3H6, Canada 5 Department of Materials Science \& Engineering, University of Toronto, Toronto, ON M5S 3E4, Canada 6 Department of Chemical Engineering \& Applied Chemistry, University of Toronto, Toronto, ON M5S 3E5, Canada 7 Canadian Institute for Advanced Research (CIFAR), Toronto, ON M5G 1M1, Canada

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23698 2025-09-30 cs.CL 57%

VIVA+: Human-Centered Situational Decision-Making

Zhe Hu, Yixiao Ren, Guanzhong Liu, Jing Li, Yu Yin

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Research Centre for Data Science & Artificial Intelligence(数据科学与人工智能研究中心) Department of Computer and Data Sciences, Case Western Reserve University(凯斯西储大学计算机与数据科学系)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21072 2025-09-26 cs.AI 57%

Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution

Kaiwen He, Zhiwei Wang, Chenyi Zhuang, Jinjie Gu

机构 * AWorld Team, Inclusion AI(Inclusion AI)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20841 2025-09-26 cs.RO cs.AI cs.LG 57%

ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation

Dekun Lu, Wei Gao, Kui Jia

机构 * Dekun Lu ∗ , Wei Gao ∗ and Kui Jia ∗(作者)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments First two authors contribute equally. Project page: https://sites.google.com/view/imaginationpolicy

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16803 2025-09-25 cs.RO cs.CV cs.NI eess.IV 57%

RG-Attn: Radian Glue Attention for Multi-modality Multi-agent Cooperative Perception

Lantao Li, Kang Yang, Wenqi Zhang, Xiaoxue Wang, Chen Sun

机构 * Sony (China) Limited(索尼(中国)有限公司) Renmin University of China(中国人民大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025 DriveX workshop (Final Version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18372 2025-09-24 cs.CV 57%

TinyBEV: Cross Modal Knowledge Distillation for Efficient Multi Task Bird's Eye View Perception and Planning

Reeshad Khan, John Gauch

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15293 2025-09-23 cs.CV cs.RO 57%

How Good are Foundation Models in Step-by-Step Embodied Reasoning?

Dinura Dissanayake, Ahmed Heakl, Omkar Thawakar, Noor Ahsan, Ritesh Thawkar, Ketan More, Jean Lahoud, Rao Anwer, Hisham Cholakkal, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学) Linköping University(林雪平大学) Australian National University(澳大利亚国立大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Project page: https://mbzuai-oryx.github.io/FoMER-Bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17328 2025-09-23 cs.CV cs.HC 57%

UIPro: Unleashing Superior Interaction Capability For GUI Agents

Hongxin Li, Jingran Su, Jingfan Chen, Zheng Ju, Yuntao Chen, Qing Li, Zhaoxiang Zhang

机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) New Laboratory of Pattern Recognition (NLPR), CASIA(中国科学院自动化所模式识别新实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(中国科学院多模态人工智能系统国家重点实验室) Hong Kong Institute of Science & Innovation, CASIA(中国科学院香港创新科学研究院) PolyU Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20148 2025-09-22 cs.AI cs.HC cs.MA 57%

The Anatomy of a Personal Health Agent

A. Ali Heydari, Ken Gu, Vidya Srinivas, Hong Yu, Zhihan Zhang, Yuwei Zhang, Akshay Paruchuri, Qian He, Hamid Palangi, Nova Hammerquist, Ahmed A. Metwally, Brent Winslow, Yubin Kim, Kumar Ayush, Yuzhe Yang, Girish Narayanswamy, Maxwell A. Xu, Jake Garrison, Amy Armento Lee, Jenny Vafeiadou, Ben Graef, Isaac R. Galatzer-Levy, Erik Schenck, Andrew Barakat, Javier Perez, Jacqueline Shreibati, John Hernandez, Anthony Z. Faranesh, Javier L. Prieto, Connor Heneghan, Yun Liu, Jiening Zhan, Mark Malhotra, Shwetak Patel, Tim Althoff, Xin Liu, Daniel McDuff, Xuhai "Orson" Xu

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Minor updates to the manuscript (V2)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14030 2025-09-18 cs.AI 57%

CrowdAgent: Multi-Agent Managed Multi-Source Annotation System

Maosheng Qin, Renyu Zhu, Mingxuan Xia, Chenkai Chen, Zhen Zhu, Minmin Lin, Junbo Zhao, Lu Xu, Changjie Fan, Runze Wu, Haobo Wang

机构 * Zhejiang University(浙江大学) NetEase Fuxi AI Lab(网易凤凰AI实验室) Zhejiang Sci-tech University(浙江科技学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10576 2025-09-16 cs.CY cs.AI 57%

Aesthetic Experience and Educational Value in Co-creating Art with Generative AI: Evidence from a Survey of Young Learners

Chengyuan Zhang, Suzhe Xu

机构 * Huaqiao University(华侨大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03747 2025-09-15 cs.SI cs.AI cs.LG 57%

Data-Driven Discovery of Mobility Periodicity for Understanding Urban Systems

Xinyu Chen, Qi Wang, Yunhan Zheng, Nina Cao, HanQin Cai, Jinhua Zhao

机构 * Massachusetts Institute of Technology(麻省理工学院) Northeastern University(东北大学) University of Central Florida(佛罗里达中央大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02734 2025-09-15 eess.IV cs.CV cs.NE stat.AP stat.ML 57%

Integrative Variational Autoencoders for Generative Modeling of an Image Outcome with Multiple Input Images

Bowen Lei, Yeseul Jeon, Rajarshi Guhaniyogi, Aaron Scheffler, Bani Mallick, Alzheimer's Disease Neuroimaging Initiatives

机构 * Alzheimer’s Disease Neuroimaging Initiatives(阿尔茨海默病神经成像计划)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03700 2025-09-12 cs.HC cs.AI 57%

MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning

Liujian Tang, Shaokang Dong, Yijia Huang, Minqi Xiang, Hongtao Ruan, Bin Wang, Shuo Li, Zhiheng Xi, Zhihui Cao, Hailiang Pang, Heng Kong, He Yang, Mingxu Chai, Zhilin Gao, Xingyu Liu, Yingnan Fu, Jiaming Liu, Xuanjing Huang, Yu-Gang Jiang, Tao Gui, Qi Zhang, Kang Wang, Yunke Zhang, Yuran Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06269 2025-09-09 cs.AI 57%

REMI: A Novel Causal Schema Memory Architecture for Personalized Lifestyle Recommendation Agents

Vishal Raman, Vijai Aravindh R, Abhijith Ragav

机构 * Radian Group Inc.(Radian集团) Sri Sivasubramaniya Nadar College Of Engineering(Sri Sivasubramaniya纳达尔工程学院) Amazon(亚马逊)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 8 pages, 2 figures, Accepted at the OARS Workshop, KDD 2025, Paper link: https://oars-workshop.github.io/papers/Raman2025.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03990 2025-09-09 cs.AI 57%

Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM Agent

Chunlong Wu, Ye Luo, Zhibo Qu, Min Wang

机构 * Tongji University(同济大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00054 2025-09-09 cs.RO cs.AI 57%

Robotic Fire Risk Detection based on Dynamic Knowledge Graph Reasoning: An LLM-Driven Approach with Graph Chain-of-Thought

Haimei Pan, Jiyun Zhang, Qinxi Wei, Xiongnan Jin, Chen Xinkai, Jie Cheng

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments We have decided to withdraw this paper as the work is still undergoing further refinement. To ensure the clarity of the results, we prefer to make additional improvements before resubmission. We appreciate the readers' understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03460 2025-09-09 cs.AI 57%

Multi-Agent Reasoning for Cardiovascular Imaging Phenotype Analysis

Weitong Zhang, Mengyun Qiao, Chengqi Zang, Steven Niederer, Paul M Matthews, Wenjia Bai, Bernhard Kainz

机构 * Department of Computing, Imperial College London, London, UK(帝国理工学院计算机系) Department of Mechanical Engineering, University College London, London, UK(伦敦大学学院机械工程系) Department of Brain Sciences, Imperial College London, London, UK(帝国理工学院脑科学系) Data Science Institute, Imperial College London, London, UK(帝国理工学院数据科学研究所) University of Tokyo, Tokyo, JP(东京大学) National Heart and Lung Institute, Imperial College London, London, UK(帝国理工学院国家心脏和肺研究所) FAU Erlangen-Nürnberg, Erlangen, DE(埃朗根-纽伦堡大学) UK Dementia Research Institute, Imperial College London, London, UK(英国痴呆研究所在伦敦帝国理工学院) Rosalind Franklin Institute, Harwell Science and Innovation Campus, Didcot, UK(罗莎琳德·弗兰克林研究所)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05338 2025-09-09 cs.RO cs.AI 57%

Plantbot: Integrating Plant and Robot through LLM Modular Agent Networks

Atsushi Masumori, Norihiro Maruyama, Itsuki Doi, johnsmith, Hiroki Sato, Takashi Ikegami

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03864 2025-09-09 cs.AI 57%

Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety

Zhenyu Pan, Yiting Zhang, Yutong Zhang, Jianshu Zhang, Haozheng Luo, Yuwei Han, Dennis Wu, Hong-Yu Chen, Philip S. Yu, Manling Li, Han Liu

机构 * Northwestern University(西北大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments accepted by the Trustworthy FMs workshop in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏