arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

2607.08193 2026-07-10 cs.LG cs.AI 新提交 74%

Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs

通过多模态大语言模型对策略进行视觉检查实现开放式多智能体自动课程

Lorenzo Pantè, Andrea Fanti, Roberto Capobianco

机构 * Sapienza University of Rome(罗马第一大学) Sony AI(索尼人工智能)

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

AI总结 研究强化学习中开放式多智能体自动课程设计难题,提出通过策略视觉检查(VIP)利用视频语言模型处理视频并提供课程建议,经星际争霸多智能体挑战赛实证,显示该方法能生成更有效课程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29399 2026-06-30 cs.AI 74%

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents

LLM引导的规划:多模态核监管文档的多跳推理

Mingyu Jeon, Bokyeong Kim, Suwan Cho, Jae Young Suh, Yonggyun Yu

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 提出LLM引导的规划方法,通过动态知识图谱状态和文档树工具,在多跳推理任务中实现81.5%准确率,显著优于无状态规划方法。

Comments Accepted at the Second Workshop on Agents in the Wild: Safety, Security, and Beyond @ ICML 2026. 8 pages (main), 3 figures, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27671 2026-06-29 cs.CV 新提交 74%

Multi-Modal Conditioned High-Resolution Transformer for Urban Electromagnetic Field Map Prediction Download PDF

面向城市电磁场地图预测的多模态条件高分辨率Transformer

Do-Eon Kim, Dongryul Park, Seungyoung Ahn, Namwoo Kang, Seong-heum Kim, Seongsin Kim

机构 * Soongsil University(崇实大学) Cho Chun Shik Graduate School of Mobility, Korea Advanced Institute of Science and Technology(韩国科学技术院赵春植移动研究生院) Department of Intelligent Semiconductors, Soongsil University(崇实大学智能半导体系)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

AI总结 提出多条件密集预测框架,利用高分辨率Transformer骨干网络,结合特征线性调制和交叉注意力机制,从建筑布局和天线配置生成500×500电磁场地图,通过复合损失函数和测试时增强提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20681 2026-06-23 cs.CV cs.LG 新提交 74%

A UAV-Based Multi-Modal Vision System for Automated Sideslope Deformation Monitoring and Hazard Detection

基于无人机多模态视觉系统的边坡变形自动监测与灾害检测

Jingfeng Zhang, Yi Li, Xianchong Liang, Huan Yang

机构 * South China Normal University(华南师范大学)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

AI总结 提出一种基于无人机LiDAR的自动化边坡灾害检测流程,包括数据采集、地面点提取、单次观测灾害筛查和多次观测变形监测,实现厘米级形变量化与危险区域识别。

Comments 29 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15872 2026-06-16 cs.CL 新提交 74%

SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks

SciOrch: 学习编排专家大语言模型以解决前沿多模态科学推理任务

Jingru Guo, Xiangyuan Xue, Lian Zhang, Wanghan Xu, Siki Chen, Philip Torr, Wanli Ouyang, Lei Bai, Zhenfei Yin

机构 * Imperial College London(伦敦帝国学院) The Chinese University of Hong Kong(香港中文大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Oxford(牛津大学) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 多模态Agent :multimodal(title);分类 cs.CL

AI总结 提出SciOrch框架,训练轻量级8B模型编排多个前沿大语言模型,通过MCTS和GRPO优化,在科学推理任务上超越最强单模型和多智能体基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03223 2026-06-03 cs.RO cs.AI 74%

BotDirector: Robot Storytelling Across the Symmetrical Reality with Multi-modal Interactions

BotDirector:跨对称现实的多模态交互机器人讲故事

Zhe Sun, Meng Wang, Lei Wang, Yuxi Wang, Wanxin Li, Yujia Peng, Zhenliang Zhang

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing, China(国家一般人工智能重点实验室,BIGAI,北京,中国) Peking University, Beijing, China(北京大学,北京,中国)

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

AI总结 提出一个结合具身交互和自然语言交互的机器人讲故事系统,利用LLM代理将儿童创建的叙事转化为自导航群体机器人的运动序列,支持灵活场景和日常物品。

Journal ref 2026 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25901 2026-05-26 cs.CV cs.RO 74%

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models

AgentGrounder:使用多模态语言模型的零样本3D视觉点云定位

Cuong Huynh, Maxim Popov, Denis Gridusov, Sergey Kolyubin

机构 * Biomechatronics and Energy-Efficient Robotics (BE2R) Lab, ITMO University(生物机械与高效能机器人实验室,ITMO大学)

专题命中 多模态Agent :multimodal(title);分类 cs.CV

AI总结 提出AgentGrounder框架,通过两阶段设计(离线构建对象查找表和在线工具驱动代理)实现零样本3D视觉定位,在ScanRefer和Nr3D上分别提升2.5%和6.3%的准确率。

Comments Code: https://github.com/be2rlab/AgentGrounder

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20734 2026-05-21 cs.CR cs.AI 74%

An Application-Layer Multi-Modal Covert-Channel Reference Monitor for LLM Agent Egress

面向LLM代理出站的应用层多模态隐通道参考监控器应用

Alfredo Metere

机构 * Enclawed, LLC(Enclawed公司)

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

AI总结 本文提出了一种应用层多模态隐通道参考监控器,用于检测和防止LLM代理在消息中泄露数据,通过多阶段文本管道、媒体加密器和残余容量测量来实现对隐通道的监控和管理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08278 2026-04-30 cs.CV cs.HC cs.RO 74%

A Multimodal Depth-Aware Method For Embodied Reference Understanding

一种多模态深度感知方法用于具身参照理解

Fevziye Irem Eyiokur, Dogucan Yaman, Hazım Kemal Ekenel, Alexander Waibel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) KIT Campus Transfer GmbH (KCT)(KIT校园转移有限公司) Istanbul Technical University(伊斯坦布尔技术大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态Agent :multimodal(title);分类 cs.CV

AI总结 本文提出一种多模态深度感知方法,结合语言模型数据增强、深度图模态和深度感知决策模块,提升复杂环境中的参照物识别准确性。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24793 2026-04-29 eess.IV cs.CV 74%

CRC-SAM: SAM-Based Multi-Modal Segmentation and Quantification of Colorectal Cancer in CT, Colonoscopy, and Histology Images

CRC-SAM:基于SAM的多模态结直肠癌在CT、结肠镜和组织学图像中的分割与量化

Daniel Lao

机构 * Independent researcher(独立研究者)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

AI总结 本文提出CRC-SAM框架,实现跨结肠镜、CT和病理图像的结直肠癌分割,通过LoRA层实现高效领域迁移,实验显示在多个数据集上优于现有方法。

Comments 4 pages, 3 figures, ISBI 2026 oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24441 2026-04-28 cs.CV 74%

AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark

AutoGUI-v2:一个全面的多模态GUI功能理解基准

Hongxin Li, Xiping Wang, Jingran Su, Zheng Ju, Yuntao Chen, Qing Li, Zhaoxiang Zhang

机构 * University of Chinese Academy of Sciences(中国科学院大学) New Laboratory of Pattern Recognition(模式识别新实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Hong Kong Institute of Science & Innovation(香港科学创新研究院) PolyU Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

AI总结 AutoGUI-v2通过多平台截图递归解析生成多样化任务,评估深度GUI功能理解和交互预测能力,揭示VLMs在功能接地与描述上的差异及复杂交互逻辑的挑战。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15327 2026-04-20 cs.HC cs.AI 74%

Eco-Bee: A Personalised Multi-Modal Agent for Advancing Student Climate Awareness and Sustainable Behaviour in Campus Ecosystems

Eco-Bee:一种面向校园生态系统的个性化多模态代理,用于提升学生气候意识和可持续行为

Caleb Adu, Neil Kapadia, Binhe Liu, Jonathan Randall, Sruthi Viswanathan

机构 * University of Hull(赫尔大学) The Spaceship Academy(太空学院) City St George’s, University of London(伦敦大学城市学院) King’s College London(伦敦国王学院) Lancaster University(兰卡斯特大学) University of Oxford(牛津大学)

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

AI总结 Eco-Bee通过整合大语言模型、行星边界框架(Eco-Score)和对话代理,为学生提供个性化反馈和行为激励,推动校园可持续发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17188 2026-04-14 cs.AI 74%

PosterGen: Aesthetic-Aware Multi-Modal Paper-to-Poster Generation via Multi-Agent LLMs

PosterGen:基于多智能体LLM的美观化论文到海报生成

Zhilin Zhang, Xiang Zhang, Jiaqi Wei, Yiwei Xu, Chenyu You

机构 * Stony Brook University(石溪大学) New York University(纽约大学) University of British Columbia(不列颠哥伦比亚大学) Zhejiang University(浙江大学) University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

AI总结 PosterGen通过多智能体LLM实现论文到海报的美观化生成,包含四个协作专业代理,提升设计质量和视觉吸引力。

Comments Project Website: https://Y-Research-SBU.github.io/PosterGen

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17312 2026-04-02 cs.CV 74%

CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning

CodeDance: 一种动态集成工具的多模态大语言模型用于可执行视觉推理

Qi Song, Honglin Li, Yingchen Yu, Haoyi Zhou, Lin Yang, Song Bai, Qi She, Zilong Huang, Yunqing Zhao

机构 * Beihang University(北京航空航天大学) Westlake University(西湖大学) ByteDance Singapore(字节跳动新加坡) ByteDance China(字节跳动中国)

专题命中 多模态Agent :MLLM(title);分类 cs.CV

AI总结 CodeDance通过可执行代码实现视觉推理,结合多工具协作与自检机制,提升复杂任务的灵活性和可解释性,实验显示其在多个基准测试中优于现有方法。

Comments CVPR 2026. Project page: https://codedance-vl.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02842 2026-03-31 cs.CL 74%

A Browser-based Open Source Assistant for Multimodal Content Verification

基于浏览器的开源多模态内容验证助手

Rosanna Milner, Michael Foster, Twin Karmakharm, Olesya Razuvayevskaya, Ian Roberts, Valentin Porcellini, Denis Teyssou, Kalina Bontcheva

机构 * University of Sheffield(谢菲尔德大学) AFP Medialab(法新社媒体实验室)

专题命中 多模态Agent :multimodal(title);分类 cs.CL

AI总结 本文提出VERIFICATION ASSISTANT,一种浏览器工具,整合多个NLP服务,为非专家用户提供清晰的可信度信号和反虚假信息指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21577 2026-03-24 cs.AI 74%

Mind over Space: Can Multimodal Large Language Models Mentally Navigate?

心灵超越空间:多模态大语言模型能否进行心理导航?

Qihui Zhu, Shouwei Ruan, Xiao Yang, Hao Jiang, Yao Huang, Shiji Zhao, Hanwei Fan, Hang Su, Xingxing Wei

机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院) Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(清华大学-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学计算机科学与技术系) School of Automation Science and Electrical Engineering , Beihang University(北京航空航天大学自动化科学与电气工程学院) college of AI, Tsinghua University(清华大学人工智能学院) Department of Computer Science and Technology , Tsinghua University(清华大学计算机科学与技术系)

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 本文提出Video2Mental基准测试,评估多模态大语言模型的空间导航能力,发现标准预训练模型无法自然生成空间表示,NavMind通过显式认知地图提升导航性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02366 2026-03-04 cs.HC cs.AI 74%

PlayWrite: A Multimodal System for AI Supported Narrative Co-Authoring Through Play in XR

PlayWrite: 一种通过XR中的游戏化互动实现AI支持的叙事协作写作的多模态系统

Esen K. Tütüncü, Qian Zhou, Frederik Brudy, George Fitzmaurice, Fraser Anderson

机构 * Institute of Neurosciences of the University of Barcelona(巴塞罗那大学神经科学研究所) Autodesk Research(Autodesk研究)

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 PlayWrite是一种通过XR中的游戏化互动实现AI支持的叙事协作写作的多模态系统,通过直接操控虚拟角色和道具,结合多智能体AI管道和大型语言模型,促进高度即兴和游戏化的创作过程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15460 2026-02-18 cs.LG cs.CV 74%

On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks

在简单视觉规划任务中多模态大语言模型推理的分布外泛化

Yannic Neuhaus, Nicolas Flammarion, Matthias Hein, Francesco Croce

机构 * Tübingen AI Center -- University of Tübingen(图宾根人工智能中心 -- 图宾根大学) EPFL(瑞士联邦理工学院) ELLIS Institute Finland -- Aalto University(芬兰ELLIS研究所 -- 阿尔托大学)

专题命中 多模态Agent :multimodal(title);分类 cs.CV

AI总结 研究多模态大语言模型在简单视觉规划任务中推理能力的分布外泛化,发现结合多种文本格式的推理方法效果最佳,纯文本模型表现优于图像输入模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01983 2026-02-03 cs.AI 74%

Evolving from Tool User to Creator via Training-Free Experience Reuse in Multimodal Reasoning

从工具使用者到创作者的进化:通过无训练经验重用在多模态推理中

Xintian Shen, Jiawei Chen, Lihao Zheng, Hao Ma, Tao Wei, Kun Zhan

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 UCT框架通过无训练经验重用,使智能体从工具使用者转变为工具创造者,提升多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12582 2026-01-21 cond-mat.mtrl-sci cs.AI 74%

Ontology-aligned structuring and reuse of multimodal materials data and workflows towards automatic reproduction

面向多模态材料数据和工作流的本体对齐结构化与重用,以实现自动重现

Sepideh Baghaee Ravari, Abril Azocar Guzman, Sarath Menon, Stefan Sandfeld, Tilmann Hickel, Markus Stricker

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 本文提出一种基于本体驱动和大型语言模型的框架,用于自动提取和结构化多模态材料数据和工作流,以提高计算结果的可重现性和重用性。

Comments 39 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17250 2025-12-22 cs.AI 74%

Accelerating Multi-modal LLM Gaming Performance via Input Prediction and Mishit Correction

通过输入预测和误位修正加速多模态大语言模型游戏性能

Ziyang Lin, Zixuan Sun, Sanhorn Chen, Xiaoyang Chen, Roy Zhao

专题命中 多模态Agent :multi-modal(title);分类 cs.AI

AI总结 通过输入预测和误位修正方法,显著降低多模态大语言模型游戏性能的推理延迟,提升整体控制效果。

Comments UIUC 25 Fall CS 498

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05335 2025-11-04 cs.CV 74%

New multimodal similarity measure for image registration via modeling local functional dependence with linear combination of learned basis functions

Joel Honkamaa, Pekka Marttinen

机构 * Department of Computer Science(计算机科学系)

专题命中 多模态Agent :multimodal(title);分类 cs.CV

Comments Improved experimental setup

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10255 2025-10-22 cs.CV cs.RO 74%

When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models

Xianzheng Ma, Brandon Smart, Yash Bhalgat, Shuai Chen, Xinghui Li, Jian Ding, Jindong Gu, Dave Zhenyu Chen, Songyou Peng, Jia-Wang Bian, Philip H Torr, Marc Pollefeys, Matthias Nießner, Ian D Reid, Angel X. Chang, Iro Laina, Victor Adrian Prisacariu

机构 * University of Oxford(牛津大学) King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学) Technical University of Munich(慕尼黑技术大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Simon Fraser University(西蒙·弗雷泽大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

Comments 2nd version update to Jun.2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05520 2025-08-27 cs.AI 74%

Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQA

Karishma Thakrar, Shreyas Basavatia, Akshay Daftardar

机构 * Georgia Institute of Technology Atlanta, GA, USA(佐治亚理工学院)

专题命中 多模态Agent :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03986 2025-08-07 cs.AI 74%

The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?

Yuan Xun, Xiaojun Jia, Xinwei Liu, Hua Zhang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Nanyang Technological University(南洋理工大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态Agent :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16940 2025-07-24 cs.CV cs.LG cs.MA 74%

AURA: A Multi-Modal Medical Agent for Understanding, Reasoning & Annotation

Nima Fathi, Amar Kumar, Tal Arbel

机构 * Center for Intelligent Machines, McGill University, Montreal, Canada(麦吉尔大学智能机器中心,加拿大蒙特利尔) Mila - Quebec AI institute, Montreal, Canada(魁北克AI研究所)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

Comments 9 pages, 3 figures, International Conference on Medical Image Computing and Computer-Assisted Intervention

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04595 2025-06-06 cs.CV 74%

Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual Learning

Ziqi Jia, Anmin Wang, Xiaoyang Qu, Xiaowen Yang, Jianzong Wang

机构 * Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司) Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

Comments Accepted by the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21347 2025-05-01 cs.AI cs.HC 74%

IRL Dittos: Embodied Multimodal AI Agent Interactions in Open Spaces

Seonghee Lee, Denae Ford, John Tang, Sasa Junuzovic, Asta Roseway, Ed Cutrell, Kori Inkpen

机构 * Stanford University(斯坦福大学) Microsoft Research(微软研究院)

专题命中 多模态Agent :multimodal(title);分类 cs.AI

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20028 2025-03-27 cs.AI 74%

OmniNova:A General Multimodal Agent Framework

Pengfei Du

专题命中 多模态Agent :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04730 2025-03-10 cs.CL cs.HC 74%

WinClick: GUI Grounding with Multimodal Large Language Models

Zheng Hui, Yinheng Li, Dan zhao, Tianyi Chen, Colby Banbury, Kazuhito Koishida

专题命中 多模态Agent :multimodal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏