arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

1905.01752 2019-05-07 cs.CV 79%

Understanding urban landuse from the above and ground perspectives: a deep learning, multimodal solution

Shivangi Srivastava, John E. Vargas-Muñoz, Devis Tuia

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Journal ref Remote Sensing of Environment, 228, pages 129 - 143, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.01560 2019-05-07 cs.AI cs.RO 79%

Dynamic Real-time Multimodal Routing with Hierarchical Hybrid Planning

Shushman Choudhury, Jacob P. Knickerbocker, Mykel J. Kochenderfer

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 8 figures, Accepted to Intelligent Vehicles (IV) Symposium 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.10763 2018-11-28 cs.CV 79%

Quality-Aware Multimodal Saliency Detection via Deep Reinforcement Learning

Xiao Wang, Tao Sun, Rui Yang, Chenglong Li, Bin Luo, Jin Tang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.09562 2018-07-26 cs.CV 79%

Change Detection between Multimodal Remote Sensing Data Using Siamese CNN

Zhenchao Zhang, George Vosselman, Markus Gerke, Devis Tuia, Michael Ying Yang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.00528 2018-04-03 cs.CV 79%

Multimodal Biometric Authentication Using Choquet Integral and Genetic Algorithm

Anouar Ben Khalifa, Sami Gazzah, Najoua Essoukri Ben Amara

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.02565 2018-02-21 cs.HC cs.AI cs.LG stat.ML 79%

Applying Cooperative Machine Learning to Speed Up the Annotation of Social Signals in Large Multi-modal Corpora

Johannes Wagner, Tobias Baur, Yue Zhang, Michel F. Valstar, Björn Schuller, Elisabeth André

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.04486 2017-10-13 cs.HC cs.CV stat.ML 79%

Multimodal Observation and Interpretation of Subjects Engaged in Problem Solving

Thomas Guntz, Raffaella Balzarini, Dominique Vaufreydaz, James L. Crowley

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Journal ref 1st Workshop on "Behavior, Emotion and Representation: Building Blocks of Interaction'', Oct 2017, Bielefeld, Germany. 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1607.04376 2016-07-18 cs.RO cs.AI 79%

Intrinsically Motivated Multimodal Structure Learning

Jay Ming Wong, Roderic A. Grupen

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1604.07806 2016-04-27 cs.AI cs.NE 79%

Using Indirect Encoding of Multiple Brains to Produce Multimodal Behavior

Jacob Schrum, Joel Lehman, Sebastian Risi

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1403.1902 2015-02-04 cs.CV 79%

Quality-based Multimodal Classification Using Tree-Structured Sparsity

Soheil Bahrampour, Asok Ray, Nasser M. Nasrabadi, Kenneth W. Jenkins

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Comments To Appear in 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2014)

Journal ref CVPR 2014, pp. 4114 - 4121

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08183 2026-08-11 cs.RO 新提交 79%

Multi-modal Interactive Control of Robotic Arm based on Offline Large Language Models

基于离线大语言模型的机械臂多模态交互控制

Hanxiao Chen

机构 * University of Tokyo(东京大学)

专题命中 多模态Agent :multi-modal(title,abstract);multimodal(comments)

AI总结 该研究提出“Socratic Models-ChatGLM”算法,基于离线开源大语言模型与PyBullet平台实现机械臂多模态交互控制,可降低成本并解决复杂多步骤机械操作任务。

Comments This research work has been accepted for poster presentation at ICRA 2026 MEI (Multimodal Embodied Interaction in Robots) Workshop. (Here is the workshop-version short paper.)

Journal ref https://ieeexplore.ieee.org/abstract/document/11371765; 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09771 2026-02-11 physics.app-ph 79%

A large scale multi-modal workflow for battery characterization: from concept to implementation

大规模多模态工作流程用于电池表征:从概念到实现

François Cadiou, Cinthya Herrera, Duncan Atkins, Elixabete Ayerbe, Giorgio Baraldi, Stéphanie Belin, Anass Benayad, Didier Blanchard, Federico Capone, Ennio Capria, Isidora Cekic Laskovic, Robert Dominko, Kristina Edström, Ajay Gautam, Lukas Helfen, Antonella Iadecola, Quentin Jacquet, Gregor Kapun, Xinyu Li, Aleksandar Matic, Nataliia Mozhzhukhina, Andrew J Naylor, Poul Norby, Chris O Keefe, Alexandre Ponrouch, Jean Pascal Rueff, Elena Tchernykova, Deyana Tchitchekova, Israel Temprano, Nikita Vostrov, Marnix Wagemaker, Martin Winter, Christian Wölke, Tejs Vegge, Sandrine Lyonnard

专题命中 多模态Agent :multi-modal(title,comments);multimodal(abstract)

AI总结 本文提出了一种大规模多模态工作流程,用于电池表征,通过整合多种技术分析电极材料的演变和电解质成分的影响,展示了标准化流程和多属性元视图的应用。

Comments 32 pages, 6 figures, research article, keywords: Battery, Workflow, Experimental characterization, Multi-modal, Large scale, Correlation, Aging

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09822 2025-07-24 cs.RO cs.SY eess.SY 79%

Active Probing with Multimodal Predictions for Motion Planning

Darshan Gadginmath, Farhad Nawaz, Minjun Sung, Faizan M Tariq, Sangjae Bae, David Isele, Fabio Pasqualetti, Jovin D'sa

机构 * Honda Research Institute, USA(本田研究院(美国)) Department of Mechanical Engineering, University of California Riverside(加州大学河滨分校机械工程系)

专题命中 多模态Agent :multimodal(title,abstract)

Comments To appear at IROS '25. 8 pages. 3 tables. 6 figures. Project page: https://darshangm.github.io/papers/active-probing-multimodal-predictions/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06136 2023-10-11 cs.HC 79%

Predicting Player Engagement in Tom Clancy's The Division 2: A Multimodal Approach via Pixels and Gamepad Actions

Kosmas Pinitas, David Renaudie, Mike Thomsen, Matthew Barthet, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis

专题命中 多模态Agent :multimodal(title,abstract)

Comments 8 pages, accepted for publication and presentation at 2023 25th ACM International Conference on Multimodal Interaction (ICMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12375 2026-07-15 cs.CV cs.AI eess.IV 新提交 79%

IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment

IQA-T1:基于工具的图像质量评估视觉证据推理

Jinjian Wu, Jiaqi Tang, Wei Wei, Yingying Yan, Jianmin Chen, Botong Geng, Lei Zhang, Qifeng Chen

机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机科学学院) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 针对开放世界图像质量评估难题,提出IQA-T1框架,通过自主调用工具生成视觉证据增强多模态大语言模型推理,构建Q-Tool数据集,实验表明该框架性能最佳且评估可解释、基于证据。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02542 2026-07-07 cs.AI cs.CV 新提交 79%

iFLYTEK-Embodied-Omni Technical Report

科大讯飞-具身全能技术报告

Yuan Zhang, Jingfei Ni, Guanchen Lu, Shiqi Zhang, Qingshan Xu, Chi Liu, Xin Nie, Wenjie Xu, Lin Gao, Zhiyuan Cheng, Mingxin Zhou, Jiajia Wu, Diyuan Liu, Jia Pan, Chao Ji

机构 * iFLYTEK LindenBot University of Science and Technology of China(中国科学技术大学)

专题命中 多模态Agent :multimodal(abstract);image-text(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

AI总结 研究通用具身智能体,提出统一多模态基础模型iFLYTEK-Embodied-Omni,通过共享多模态自注意力联合建模视觉、语言和动作,结合多种数据构建数据集并采用四阶段策略训练,实现脑-小脑协作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29445 2026-06-30 cs.CV cs.AI 79%

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

通过通用关键帧提取桥接VideoQA和视频引导的智能体任务

Sunqi Fan, Qingle Liu, Runqi Yin, Meng-Hao Guo, Shuojin Yang

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出VG-GUIBench基准和TASKER关键帧提取算法,联合考虑任务相关性和场景动态,在VideoQA和视频引导的GUI任务上提升性能。

Comments Accepted by ECCV 2026. Project Page: https://vg-gui-tasker.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26552 2026-06-26 cs.CV cs.AI 新提交 79%

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

感知、判断与进化:基于事后洞察的自优化取证智能体用于AI生成图像检测

Yangjun Wu, Keyu Yan, Yu Liu, Jingren Zhou, Fei Huang, Rong Zhang, Zhou Zhao, Fei Wu

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出ForeAgent框架,采用感知-判断架构融合多视图线索,并引入事后洞察驱动的自优化策略,通过采样-反思-进化范式持续提升检测能力,在多个基准上达到最优性能。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20717 2026-06-23 cs.CV cs.AI cs.CR 新提交 79%

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents

MIRAGE: 针对Web代理的隐蔽视觉提示注入漏洞检测

Xuelong Dai, Jianyu Ma, Boyang Ma, Biwei Yan, Yijun Yang, Yue Zhang

机构 * SDU(山东大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出MIRAGE框架,利用扩散模型在受限区域生成视觉上无害的对抗图像,实现针对多模态大模型Web代理的隐蔽间接提示注入攻击,以检测其视觉漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03005 2026-06-03 cs.CV cs.AI 79%

MUSE: A Unified Agentic Harness for MLLMs

MUSE: 多模态大语言模型的统一智能体框架

Jianglin Lu, Hailing Wang, Xu Ma, Qihua Dong, Mingyuan Zhang, Yizhou Wang, Yun Fu

机构 * Northeastern University(东北大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出MUSE框架,通过可组合模块(任务表示、视觉处理、感知工具、结构化解析、确定性验证和验证器引导修复)提升冻结多模态大语言模型性能,无需重新训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30639 2026-06-01 cs.CV cs.AI cs.RO 79%

PInVerify: An Offline Embodied Benchmark for Active Instance Verification

PInVerify:面向主动实例验证的离线具身基准

Yuhang Jiang

机构 * University of Trento(特伦托大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出主动实例验证任务,构建离线具身基准PInVerify,通过多视角导航和细粒度属性匹配评估具身智能体,并基于多模态大语言模型建立基线。

Comments Accepted as a poster at the Foundation Models Meet Embodied Agents (FMEA) Workshop, CVPR 2026. 44 pages including appendix. Code: https://github.com/Avalon-S/PInVerify

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24023 2026-01-01 cs.CV cs.AI 79%

RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations

RSAgent: 通过多轮工具调用学习推理与行动进行文本引导的分割

Xingqi He, Yujie Zhang, Shuyong Gao, Wenjie Li, Lingyi Hong, Mingxi Chen, Kaixun Jiang, Jiyuan Fu, Wenqiang Zhang

机构 * Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University(上海智能信息处理关键实验室,计算机科学与人工智能学院,复旦大学) Artificial Intelligence, Fudan University(人工智能,复旦大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 RSAgent通过多轮工具调用实现文本引导分割的推理与行动,采用两阶段框架提升分割性能,达到领域内和领域外基准的最先进水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18171 2026-08-11 cs.LG 版本更新 78%

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

FlashRT:用于引导智能体部署实时多模态应用的智能体框架

Krish Agarwal, Zhuoming Chen, Yanyuan Qin, Zhenyu Gu, Atri Rudra, Beidi Chen

机构 * Carnegie Mellon University(卡内基梅隆大学) AMD(超威半导体公司) University at Buffalo(纽约州立大学水牛城分校)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 研究实时多模态应用部署难题,提出FlashRT智能体框架。通过新范式引导编码智能体多阶段转换,将参考实现转为高效部署,在不同GPU上显著提升性能,尤其在专家优化不成熟平台更具扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00874 2026-08-11 cs.HC 版本更新 78%

MIA: A Visual Analytics System for Multimodal Spectral Imaging Data

MIA:一种用于多模态光谱成像数据的可视分析系统

Hennes Rave, Katharina Kronenberg, Hannes Gödde, Lea Tobergte, Michael Holtkamp, Julia Werner, Peter Bohrer, Fabian Lohöfer, Rickmer Braren, David Clases, Uwe Karst, Lars Linsen

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 提出一种模态无关的可视分析系统MIA,集成光谱预处理、降维、交互分割和相似性分析等全流程,支持层次化嵌入和多模态数据联合分析,通过三个实际案例验证其有效性。

Journal ref Computers & Graphics, Vol. 139, Art. 104727, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10055 2026-08-10 cs.RO 版本更新 78%

STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations

STRONG-VLA: 多模态扰动下视觉-语言-动作模型的解耦鲁棒性学习

Yuhan Xie, Yuping Yan, Yunqi Zhao, Handing Wang, Yaochu Jin

机构 * Zhejiang University(浙江大学) Westlake University(西湖大学) Xidian University(西安电子科技大学)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出STRONG-VLA框架,通过解耦鲁棒性学习与任务对齐优化,提升多模态扰动下的视觉-语言-动作模型鲁棒性与任务执行精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05966 2026-08-07 cs.HC 新提交 78%

A Modular Workflow for Multimodal Reading Experiments

多模态阅读实验的模块化工作流程

Thomas Krämer, Thomas Kosch, Dagmar Kern, Daniel Hienert

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 该研究提出一种基于网络的模块化工作流程,整合眼动、EEG等多模态数据,可适配不同实验设置,用于在线阅读的多模态研究,支持生态情境下的实证验证。

Comments In Proceedings of the 30th International Conference on Knowledge-Based and Intelligent Information & Engineering Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22014 2026-08-06 cs.MA cs.RO 版本更新 78%

DM$^3$-Nav: Decentralized Multi-Agent Multimodal Multi-Object Semantic Navigation

DM$^3$-Nav:去中心化多智能体多模态多目标语义导航

Amin Kashiri, Atharva Jamsandekar, Yasin Yazıcıoğlu

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Northeastern University(东北大学)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 DM$^3$-Nav是一种去中心化的多智能体语义导航系统,支持多模态开放词汇目标指定和多目标任务,通过局部通信实现自主协调,验证了在真实办公环境中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04398 2026-07-30 cs.HC cs.CY 版本更新 78%

The Agency Gap in AI-Supported Writing: How Reactive and Proactive Agent Designs Shape Multimodal Reasoning

AI辅助写作中的主体差距:反应式与主动式智能体设计如何塑造多模态推理

Yueqiao Jin, Kaixun Yang, Roberto Martinez-Maldonado, Dragan Gašević, Lixiang Yan

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 该研究针对AI辅助写作中的主体差距,对比反应式与主动式智能体设计的效果,发现主动式设计可强化多模态推理关联,还能缩小AI素养相关的表现差异,为教育AI智能体设计提供了方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21113 2026-07-24 cs.RO 新提交 78%

RL-MACRO: A Cybernetic Closed-Loop Intelligence Framework for Multimodal Adaptive Robotic Craniotomy

RL-MACRO:一种用于多模态自适应机器人开颅手术的控制论闭环智能框架

Xiao Zhang, Jiaxuan Li, Renzhen Le, Di Wu, Chao Sun, Jiachen Zhu, Haoyuan Zhang, Xiang Li, Jian Liu, Zhenzhi Ying, Pengfei Zhang, Liming Shu

机构 * Dalian University of Technology(大连理工大学) Second Hospital of Dalian Medical University(大连医科大学附属第二医院) The University of Tokyo(东京大学)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 研究针对自主机器人开颅手术面临的挑战,提出RL-MACRO框架,通过多模态感知、自适应决策和机器人执行实现闭环控制,利用CNN-LSTM观测器重建温度,经IQL策略和双头执行器优化行为,实验验证了该框架在骨切割方面的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19568 2026-07-23 eess.SY cs.SY 新提交 78%

From P&ID Drawings to Process Graphs: A Multimodal Language Model Approach

从管道和仪表图到过程图:一种多模态语言模型方法

Baikai Zhu, Samuel Duong, Javal Vyas, Mehmet Mercangöz

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 研究针对P&IDs数字化难题,提出基于多模态大语言模型的两阶段工作流程,将其数字化重定义为设备标签提取与拓扑推理,经案例研究评估,该方法比端到端数字化更优,凸显知识引导工作流程在P&ID数字化中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏