arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

2412.17288 2024-12-24 cs.RO cs.AI 74%

Multi-Modal Grounded Planning and Efficient Replanning For Learning Embodied Agents with A Few Examples

Taewoong Kim, Byeonghwi Kim, Jonghyun Choi

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

Comments AAAI 2025 (Project page: https://twoongg.github.io/projects/flare/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00627 2024-12-10 cs.HC cs.AI 74%

ARChef: An iOS-Based Augmented Reality Cooking Assistant Powered by Multimodal Gemini LLM

Rithik Vir, Parsa Madinei

专题命中 多模态Agent :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09971 2024-11-18 cs.CV cs.RO 74%

Explanation for Trajectory Planning using Multi-modal Large Language Model for Autonomous Driving

Shota Yamazaki, Chenyu Zhang, Takuya Nanri, Akio Shigekane, Siyuan Wang, Jo Nishiyama, Tao Chu, Kohei Yokosawa

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

Comments Accepted and presented at ECCV 2024 2nd Workshop on Vision-Centric Autonomous Driving (VCAD) on September 30, 2024. 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20409 2024-11-01 cs.CV physics.med-ph 74%

Physics-Regularized Multi-Modal Image Assimilation for Brain Tumor Localization

Michal Balcerak, Tamaz Amiranashvili, Andreas Wagner, Jonas Weidner, Petr Karnakov, Johannes C. Paetzold, Ivan Ezhov, Petros Koumoutsakos, Benedikt Wiestler, Bjoern Menze

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11347 2024-09-18 cs.AI 74%

Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments

Takanori Ugai, Kensho Hara, Shusaku Egami, Ken Fukuda

专题命中 多模态Agent :multimodal(title);分类 cs.AI

Comments 5 pages, 1 figure, 1 table, accepted in Embodied AI 2024 Workshop held in conjunction with CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06978 2024-09-02 cs.CV 74%

Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation

Zhenxin Li, Kailin Li, Shihao Wang, Shiyi Lan, Zhiding Yu, Yishen Ji, Zhiqi Li, Ziyue Zhu, Jan Kautz, Zuxuan Wu, Yu-Gang Jiang, Jose M. Alvarez

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments The 1st place solution of End-to-end Driving at Scale at the CVPR 2024 Autonomous Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06720 2024-08-26 cs.CV cs.LG q-bio.QM 74%

Multimodal Analysis of White Blood Cell Differentiation in Acute Myeloid Leukemia Patients using a β-Variational Autoencoder

Gizem Mert, Ario Sadafi, Raheleh Salehi, Nassir Navab, Carsten Marr

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments Accepted for publication at MICCAI 2024 workshop on AI for Imaging Genomics Learning (AIIG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00535 2024-07-02 cs.CE cs.CV 74%

AI-powered multimodal modeling of personalized hemodynamics in aortic stenosis

Caglar Ozturk, Daniel H. Pak, Luca Rosalia, Debkalpa Goswami, Mary E. Robakowski, Raymond McKay, Christopher T. Nguyen, James S. Duncan, Ellen T. Roche

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments CO and DHP contributed equally to this work. JSD and ETR are corresponding authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.05295 2024-06-19 cs.AI cs.FL 74%

Multimodal Pretrained Models for Verifiable Sequential Decision-Making: Planning, Grounding, and Perception

Yunhao Yang, Cyrus Neary, Ufuk Topcu

专题命中 多模态Agent :multimodal(title);分类 cs.AI

Comments Accepted as full paper in AAMAS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16356 2024-04-26 cs.NI cs.AI cs.LG 74%

Integration of Mixture of Experts and Multimodal Generative AI in Internet of Vehicles: A Survey

Minrui Xu, Dusit Niyato, Jiawen Kang, Zehui Xiong, Abbas Jamalipour, Yuguang Fang, Dong In Kim, Xuemin, Shen

专题命中 多模态Agent :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.05821 2023-04-28 cs.LG cs.AI cs.DB 74%

PEg TRAnsfer Workflow recognition challenge report: Does multi-modal data improve recognition?

Arnaud Huaulmé, Kanako Harada, Quang-Minh Nguyen, Bogyu Park, Seungbum Hong, Min-Kook Choi, Michael Peven, Yunshuang Li, Yonghao Long, Qi Dou, Satyadwyoom Kumar, Seenivasan Lalithkumar, Ren Hongliang, Hiroki Matsuzaki, Yuto Ishikawa, Yuriko Harai, Satoshi Kondo, Mamoru Mitsuishi, Pierre Jannin

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

Comments Challenge report doi.org/10.1016/j.cmpb.2023.107561

Journal ref Computer Methods and Programs in Biomedicine, Volume 236, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03996 2022-05-27 cs.LG cs.CV cs.RO 74%

Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers

Ruihan Yang, Minghao Zhang, Nicklas Hansen, Huazhe Xu, Xiaolong Wang

专题命中 多模态Agent :cross-modal(title);分类 cs.CV

Comments Our project page with videos is at https://RchalYang.github.io/LocoTransformer

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15054 2021-10-29 cs.HC cs.CL cs.CY 74%

Adaptive Multimodal and Multisensory Empathic Technologies for Enhanced Human Communication

Roxana Girju

专题命中 多模态Agent :multimodal(title);分类 cs.CL

Comments 10 pages; This position paper was presented at the Rethinking the Senses: A Workshop on Multisensory Embodied Experiences and Disability Interactions associated with the ACM CHI Conference on Human Factors in Computing Systems, May 2021

Journal ref ACM CHI Conference on Human Factors in Computing Systems, May 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.14156 2021-04-30 cs.RO cs.CV 74%

Radar-based Automotive Localization using Landmarks in a Multimodal Sensor Graph-based Approach

Stefan Jürgens, Niklas Koch, Marc-Michael Meinecke

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Journal ref Proceedings of 21st International Radar Symposium (IRS 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.10384 2021-01-27 cs.RO cs.AI 74%

droidlet: modular, heterogenous, multi-modal agents

Anurag Pratik, Soumith Chintala, Kavya Srinet, Dhiraj Gandhi, Rebecca Qian, Yuxuan Sun, Ryan Drew, Sara Elkafrawy, Anoushka Tiwari, Tucker Hart, Mary Williamson, Abhinav Gupta, Arthur Szlam

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.03199 2020-10-27 cs.CV 74%

Multimodal End-to-End Autonomous Driving

Yi Xiao, Felipe Codevilla, Akhil Gurram, Onay Urfalioglu, Antonio M. López

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments The paper has been accepted by IEEE Transactions on Intelligent Transportation Systems 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.05205 2020-08-11 cs.RO cs.AI cs.MA cs.SY eess.SY 74%

Implicit Multiagent Coordination at Unsignalized Intersections via Multimodal Inference Enabled by Topological Braids

Christoforos Mavrogiannis, Jonathan A. DeCastro, Siddhartha S. Srinivasa

专题命中 多模态Agent :multimodal(title);分类 cs.AI

Comments 16 pages, 13 figures, new experiments, new explanatory figures for intuition and new title

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.03733 2020-02-11 cs.CV 74%

Robust Multimodal Image Registration Using Deep Recurrent Reinforcement Learning

Shanhui Sun, Jing Hu, Mingqing Yao, Jinrong Hu, Xiaodong Yang, Qi Song, Xi Wu

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Journal ref Asian Conference on Computer Vision (ACCV). 2018. 511-526

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.00525 2019-05-03 cs.CV 74%

3D BAT: A Semi-Automatic, Web-based 3D Annotation Toolbox for Full-Surround, Multi-Modal Data Streams

Walter Zimmer, Akshay Rangesh, Mohan Trivedi

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.03384 2019-04-04 cs.CV 74%

Prediction of laparoscopic procedure duration using unlabeled, multimodal sensor data

Sebastian Bodenstedt, Martin Wagner, Lars Mündermann, Hannes Kenngott, Beat Müller-Stich, Michael Breucha, Sören Torge Mees, Jürgen Weitz, Stefanie Speidel

专题命中 多模态Agent :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.07502 2019-02-01 cs.CV 74%

A Multimodal, Full-Surround Vehicular Testbed for Naturalistic Studies and Benchmarking: Design, Calibration and Deployment

Akshay Rangesh, Kevan Yuen, Ravi Kumar Satzoda, Rakesh Nattoji Rajaram, Pujitha Gunaratne, Mohan M. Trivedi

专题命中 多模态Agent :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17456 2026-06-17 cs.RO q-bio.NC 新提交 73%

Embodiment Shapes Rolling Behavior in a Multimodal Infant Model

具身形态塑造多模态婴儿模型中的翻滚行为

Leon Philipp, Francisco M. López, Jochen Triesch

机构 * Frankfurt Institute for Advanced Studies(法兰克福高等研究院) Goethe University Frankfurt(法兰克福大学) University of New South Wales(新南威尔士大学)

专题命中 多模态Agent :multimodal(title,comments)

AI总结 通过虚拟婴儿MIMo学习仰卧到俯卧翻滚,研究婴儿运动发展中的具身形态变化如何影响行为,发现与真实婴儿一致的发育趋势和协调模式。

Comments 7 pages, 7 figures. Accepted at the 2026 IEEE ICDL Conference. Cite as: L. Philipp, F. M. López, and J. Triesch, "Embodiment Shapes Rolling Behavior in a Multimodal Infant Model", in 2026 IEEE International Conference on Development and Learning (ICDL). IEEE, 2026, pp. 1-7

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26779 2026-06-02 cs.CV cs.AI 73%

Limits of Spatial Imagery Reasoning in Frontier LLM Models

前沿大语言模型在空间意象推理中的局限性

Sergio Y. Hayashi, Nina S. T. Hirata

机构 * Institute of Mathematics and Statistics – University of São Paulo(数学统计研究所 – 圣保罗大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 本研究通过引入外部“意象模块”辅助3D模型旋转任务,发现即使外包整体3D状态维护,前沿模型仍缺乏基础视觉空间原语,导致准确率最高仅62.5%。

Comments 25 pages. v2: Title updated; added a section on object/spatial imagery and propositional reasoning; added new experimental results for the single-object rotation probe

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30318 2026-05-29 cs.GR cs.AI cs.CV 73%

Before the Shutter: Aesthetic and Actionable Portrait Photography Planning in 3D Scenes

快门之前:3D场景中美学的且可执行的人像摄影规划

Ruixiang Jiang, Chang Wen Chen

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 提出在3D场景中生成人像姿态、相机、照明和曝光方案的方法,通过构建摄影场景图实现美学引导的规划,生成视觉上引人注目且几何与光度可行的人像。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02900 2026-05-26 cs.CR cs.AI cs.CV cs.RO 73%

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

具身人工智能的安全性:风险、攻击与防御综述

Xiao Li, Xiang Zheng, Yifeng Gao, Xinyu Xia, Yixu Wang, Xin Wang, Ye Sun, Yunhan Zhao, Ming Wen, Jiayu Li, Zixing Chen, Xun Gong, Yi Liu, Yige Li, Yutao Wu, Cong Wang, Jun Sun, Yixin Cao, Zhineng Chen, Jingjing Chen, Tao Gui, Qi Zhang, Zuxuan Wu, Xipeng Qiu, Xuanjing Huang, Tiehua Zhang, Zhipeng Wei, Kun Wang, Xinfeng Li, Hanxun Huang, Sarah Erfani, James Bailey, Jianping Wang, Chaowei Xiao, Ran He, Bo Li, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) City University of Hong Kong(香港城市大学) Jilin University(吉林大学) Singapore Management University(新加坡管理大学) Deakin University(德肯大学) Tongji University(同济大学) Nanyang Technological University(南洋理工大学) Chinese Academy of Sciences(中国科学院) The University of Melbourne(墨尔本大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

AI总结 本文综述了具身AI在感知、认知、规划、行动及交互全流程中的安全风险、攻击与防御方法,提出了多层次分类体系,并指出了多模态感知融合脆弱性、规划不稳定及人机交互可信度等关键挑战。

Comments Survey paper; 75 pages, 4 figures, 18 tables; v2 expands embodied-specific coverage of agentic threats, World Action Model threats, and contextual risk mitigation, with over 100 new references added. Project page: https://x-zheng16.github.io/Awesome-Embodied-AI-Safety/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05014 2026-04-08 cs.RO cs.AI cs.CV 73%

StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

StarVLA:一种积木式代码库,用于视觉-语言-动作模型开发

StarVLA Community

机构 * Von Neumann Institute, HKUST(香港科技大学冯·诺依曼研究所)

专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

AI总结 StarVLA通过模块化架构、可重用训练策略和统一评估接口,解决VLA方法碎片化问题,提升可复现性和跨架构兼容性。

Comments Open-source VLA infra, Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01221 2026-04-02 cs.AI cs.CV 73%

HippoCamp: Benchmarking Contextual Agents on Personal Computers

HippoCamp:在个人电脑上对上下文代理进行基准测试

Zhe Yang, Shulin Tian, Kairui Hu, Shuai Liu, Hoang-Nhat Nguyen, Yichi Zhang, Zujin Guo, Mengying Yu, Zinan Zhang, Jingkang Yang, Chen Change Loy, Ziwei Liu

机构 * S-Lab, Nanyang Technological University, Singapore(新加坡南洋理工大学S-Lab)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 HippoCamp是一个新的基准,用于评估代理在多模态文件管理中的能力,通过用户为中心的环境建模个体用户档案并搜索大规模个人文件进行上下文感知推理,揭示了当前代理在真实环境中的局限性。

Comments Project Page: https://hippocamp-ai.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14099 2026-03-16 cs.CV cs.AI 73%

FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration

FAPE-IR:面向全场景图像修复的频率感知规划与执行框架

Jingren Liu, Shuning Xu, Qirui Yang, Yun Wang, Xiangyu Chen, Zhong Ji

机构 * Tianjin University(天津大学) University of Macau(澳门大学) City University of Hong Kong(香港城市大学) Institute of Artificial Intelligence (TeleAI)(人工智能研究所(TeleAI))

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出FAPE-IR框架,通过频率感知规划与执行模块,结合多模态大语言模型和LoRA-MoE架构,实现统一且可解释的全场景图像修复,实验显示其在七项任务中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20230 2026-03-09 cs.AI cs.CV cs.MA 73%

A Multi-Agent System Enables Versatile Information Extraction from the Chemical Literature

多智能体系统实现从化学文献中灵活的信息提取

Yufan Chen, Ching Ting Leung, Bowen Yu, Jianwei Sun, Yong Huang, Linyan Li, Hao Chen, Hanyu Gao

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出了一种基于多模态大语言模型的多智能体系统,实现了从化学文献中高效提取化学信息,F1分数达76.27%,显著提升信息提取效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03198 2026-03-04 cs.RO cs.CL cs.CV 73%

ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments

ACE-Brain-0:空间智能作为通用具身化体系的共享框架

Ziyang Gong, Zehang Luo, Anke Tang, Zhe Liu, Shi Fu, Zhi Hou, Ganlin Yang, Weiyun Wang, Xiaofeng Wang, Jianbo Liu, Gen Luo, Haolan Kang, Shuang Luo, Yue Zhou, Yong Luo, Li Shen, Xiaosong Jia, Yao Mu, Xue Yang, Chunxiao Liu, Junchi Yan, Hengshuang Zhao, Dacheng Tao, Xiaogang Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) University of Science(科学技术大学) Fudan University(复旦大学) Xiamen University(厦门大学) East China Normal University(华东师范大学) Wuhan University(武汉大学) Sun Yat-sen University(中山大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

AI总结 ACE-Brain-0 通过空间智能作为共享框架,统一了自动驾驶、机器人和 UAVs 的具身化任务,采用 SSR 范式和 GRPO 方法实现跨领域泛化和领域精通的平衡。

Comments Code: https://github.com/ACE-BRAIN-Team/ACE-Brain-0 Hugging Face: https://huggingface.co/ACE-Brain/ACE-Brain-0-8B

详情

展开后加载摘要…

URL PDF HTML 收藏