arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

1804.05644 2018-04-17 cs.DS 78%

Multimodal Dynamic Journey Planning

Kalliopi Giannakopoulou, Andreas Paraskevopoulos, Christos Zaroliagis

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.09483 2017-10-27 cs.RO cs.LG 78%

Multimodal Probabilistic Model-Based Planning for Human-Robot Interaction

Edward Schmerling, Karen Leung, Wolf Vollprecht, Marco Pavone

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.06333 2017-08-22 cs.OH 78%

SigViewer: Visualizing Multimodal Signals Stored in XDF (Extensible Data Format) Files

Yida Lin, Clemens Brunner, Paul Sajda, Josef Faller

专题命中 多模态Agent :multimodal(title,abstract)

Comments 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.00470 2017-08-09 stat.ML cs.LG 78%

Learning Multimodal Transition Dynamics for Model-Based Reinforcement Learning

Thomas M. Moerland, Joost Broekens, Catholijn M. Jonker

专题命中 多模态Agent :multimodal(title,abstract)

Comments Scaling Up Reinforcement Learning (SURL) Workshop @ European Machine Learning Conference (ECML)

详情

展开后加载摘要…

URL PDF HTML 收藏
1504.03855 2015-04-17 q-bio.NC physics.med-ph 78%

A versatile clearing agent for multi-modal brain imaging

Irene Costantini, Jean-Pierre Ghobril, Antonino Paolo Di Giovanna, Anna Letizia Allegra Mascaro, Ludovico Silvestri, Marie Caroline Müllenbroich, Leonardo Onofri, Valerio Conti, Francesco Vanzi, Leonardo Sacconi, Renzo Guerrini, Henry Markram, Giulio Iannello, Francesco Saverio Pavone

专题命中 多模态Agent :multi-modal(title);multimodal(abstract)

Comments in Scientific Reports 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1303.4503 2013-03-20 physics.ins-det physics.med-ph 78%

EndoTOFPET-US a Novel Multimodal Tool for Endoscopy and Positron Emission Tomography

Erika Garutti

专题命中 多模态Agent :multimodal(title);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
0809.1074 2009-12-01 math.DS 78%

Multifractal analysis for multimodal maps

Mike Todd

专题命中 多模态Agent :multimodal(title,abstract)

Comments Minor rewrites

详情

展开后加载摘要…

URL PDF HTML 收藏
0711.2531 2009-12-01 q-bio.PE 78%

Multimodal pattern formation in phenotype distributions of sexual populations

Michael Doebeli, Hendrik J. Blok, Olof Leimar, Ulf Dieckmann

专题命中 多模态Agent :multimodal(title,abstract)

Journal ref Proc. R. Soc. B (2007) 274, 347-357

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07246 2024-10-08 cs.CL cs.AI 77%

Propaganda to Hate: A Multimodal Analysis of Arabic Memes with Multi-Agent LLMs

Firoj Alam, Md. Rafiul Biswas, Uzair Shah, Wajdi Zaghouani, Georgios Mikros

专题命中 多模态Agent :multimodal(title,comments);分类 cs.CL、cs.AI

Comments propaganda, hate-speech, disinformation, misinformation, fake news, LLMs, GPT-4, multimodality, multimodal LLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14098 2026-08-13 cs.CV 版本更新 77%

ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization

ForgeryVCR: 通过高效的取证工具在MLLMs中实现视觉中心推理用于图像伪造检测与定位

Youqi Wang, Shen Chen, Haowei Wang, Rongxuan Peng, Taiping Yao, Shunquan Tan, Changsheng Chen, Bin Li, Shouhong Ding

机构 * Shenzhen University(深圳大学) Tencent Youtu Lab(腾讯优图实验室)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 ForgeryVCR通过高效的取证工具实现视觉中心推理,提升图像伪造检测与定位的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09654 2026-08-11 cs.AI 新提交 77%

Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching

无幻觉的GUI定位:基于无回归的布局感知匹配

Yuke Li, Xuehan Hou

机构 * School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 该研究提出无回归框架,通过解耦指令理解与布局感知定位,在ScreenSpot-Pro和Mind2Web数据集上显著提升GUI定位的准确率、成功率及元素选择率,抑制了坐标幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04425 2026-08-11 cs.CL cs.AI cs.CV cs.LG cs.MM 版本更新 77%

UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents

UI-MOPD:用于持续GUI智能体学习的多平台策略蒸馏

Niu Lian, Tongbo Chen, Zhehao Yu, Chengzhen Duan, Fazhan Liu, Hui Liu, Pei Fu, Jian Luan, Heng Qu, Shu-Tao Xia, Jinpeng Wang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Xiaomi(小米) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Zhejiang University(浙江大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 针对构建多平台GUI智能体的挑战,构建高质量数据集Uni-GUI,提出UI-MOPD方法,通过多教师策略蒸馏实现持续学习,动态选教师并转移行为先验,实验证明其能平衡跨平台能力保留与新平台适应。

Comments Technical report. 27 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00066 2026-08-04 cs.CV 新提交 77%

PhysAgent: A Multi-Agent Framework for Reliable Remote Heart Rate Estimation

PhysAgent:用于可靠远程心率估计的多智能体框架

Yehui Yang, Bo Zhao, Junzhe Cao, Hui Ma, Yue Sun, Wenjin Wang, Zitong Yu

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 PhysAgent是一种推理时多智能体候选验证框架,以多个基础rPPG估计器输出为待验证生理假设,结合Qwen3-VL-4B多模态大语言模型推理与确定性融合,提升远程心率估计的稳定性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28595 2026-07-31 cs.CV 新提交 77%

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

Beacon:智能体何时及如何执行智能体视觉推理

Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen, Tianrun Xu, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Xianghua Ying

机构 * Peking University(北京大学) Kling Team(Kling团队) HKUST(GZ)(香港科技大学(广州)) CUHK(香港中文大学) ZJU(浙江大学) THU(清华大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 本研究针对现有智能体视觉推理模型模式适应性有限、工具增益被损害抵消的问题,提出Beacon模型,通过强化学习相关机制提升性能与适应性,在多基准上表现优异。

Comments 33 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25993 2026-07-29 cs.CV 新提交 77%

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

超越缩放:学习用于超高分辨率遥感的多工具视觉推理

Fengxiang Wang, Jiangnan Huang, Mingshuo Chen, Yueying Li, Yang Shi, Junwei Luo, Haoyu Wang, Yansheng Li, Jing Zhang, Haiyan Zhao, Wenjing Yang

机构 * National University of Defense Technology(国防科技大学) Wuhan University(武汉大学) Tsinghua University(清华大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 针对超高分辨率遥感图像给多模态大语言模型带来的挑战,提出GeoMTVR数据集,结合监督微调与以工具注意力为重点的强化学习算法开发GeoLens,实验证明其在多工具视觉推理方面优于直接推理和单工具放大基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28971 2026-06-30 cs.CV 77%

Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution

自进化智能体图像恢复:通过深思熟虑的规划与直觉执行

Shuang Cui, Fan Ji, Guanglong Sun, Yufei Guo, Xiongxin Tang, Jiangmeng Li, Fanjiang Xu

机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学) School of Life Sciences, Tsinghua University(清华大学生命科学学院) Intelligent Science & Technology Academy of CASIC(中国科学院 CASIC 智能科学与技术学院)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出SEAR框架,将图像恢复建模为序列决策问题,采用直觉执行器与深思熟虑规划器,结合剪枝感知蒙特卡洛树搜索和自进化情景记忆,解决贪婪搜索和信息利用不足问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26122 2026-06-26 cs.CV 新提交 77%

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents

DocArena:将原始文档转化为可控的训练环境用于文档搜索代理

Jiamian Wang, Ruiyi Zhang, Tong Yu, Jing Shi, Samyadeep Basu, Rajiv Jain, Zhiqiang Tao, Tong Sun

机构 * Rochester Institute of Technology(罗切斯特理工学院) Adobe Research(Adobe研究院)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出DocArena自动化流程,通过多模态文档结构化、推理型QA对构建和质量控制,生成可控训练环境,使基于文本LLM的搜索代理在多模态文档检索和问答中取得最佳性能。

Comments search agent for documents

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22617 2026-06-23 cs.CV 新提交 77%

OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

OmniSpace: 自动驾驶多模态大语言模型的高效几何感知

Hao Vo, Phu Loc Nguyen, Khoa Vo, Sieu Tran, Duc Minh Nguyen, Ngo Xuan Cuong, Nghi D. Q. Bui, Anh Nguyen, Duy Minh Ho Nguyen, Ngan Le

机构 * University of Arkansas(阿肯色大学) Google Research, Google(谷歌研究院) University of Liverpool(利物浦大学) Max Planck Research School for Intelligent Systems(马克斯·普朗克智能系统研究所)

专题命中 多模态Agent :MLLM(summary_cn);multimodal(abstract);分类 cs.CV

AI总结 提出OmniSpace,一种即插即用的几何感知范式,通过相机位姿注入器、多视图极线注意力模块和3D几何蒸馏目标,从纯2D观测中提升MLLM的空间推理能力,在多个自动驾驶基准上超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09110 2026-06-09 cs.CV 新提交 77%

HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging

HDRAgent: 一种用于多曝光HDR成像的智能体框架

Weiyu Zhou, Tao Hu, Yijian Wang, Xiaogang Xu, Ruixing Wang, Qingsen Yan

机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院) Shenzhen Research Institute, Northwestern Polytechnical University(西北工业大学深圳研究院) Zhejiang University(浙江大学) Camera Group, DJI(大疆相机部门)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出首个智能体驱动的HDR成像框架HDRAgent,通过细粒度上下文知识匹配、感知-失真反馈机制和智能体引导的生成对齐策略,自适应选择重建策略,减少复杂动态场景中的鬼影和局部伪影。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07549 2026-06-09 cs.AI cs.MA 新提交 77%

PathoSage: Towards Multi-Source Evidence Adjudication in Pathology via Experience-Aware Agentic Workflow

PathoSage:通过经验感知的代理工作流实现病理学多源证据裁决

Chengyang Zhang, Wenchuan Zhang, Bo Li, Mengran Li, Bob Zhang, Yuhao Yi, Hong Bu, Jiancheng Lv

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) Department of Pathology and Institute of Clinical Pathology, West China Hospital, Sichuan University(四川大学华西医院病理科/临床病理研究所) Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系) School of Intelligent Systems Engineering, Sun Yat-sen University(中山大学智能工程学院)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 提出PathoSage框架,通过结构化证据审议和Beta-Bernoulli经验系统,独立评估工具证据并解决冲突,减少幻觉和分类器分歧,提升病理学推理鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22208 2026-06-09 cs.CV 版本更新 77%

EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

EvoIR-Agent: 通过经验驱动学习实现自进化图像修复智能体

Kailin Zhuang, Jiawei Wu, Zhi Jin

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 本文提出EvoIR-Agent,通过经验驱动学习解决图像修复中经验不足导致的规划失败问题,通过构建分层经验池和自进化机制提升修复性能和效率,实验表明其在全参考指标上表现优异,且在性能与效率之间取得显著平衡。

Comments Temporarily withdrawn for institutional clearance and compliance review. A revised version will be uploaded once the process is finalized

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31174 2026-06-01 cs.CV cs.LG 77%

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning

任意场景检测:一种具有经验感知推理的目标检测智能体框架

Wenlun Zhang, Jun Yin, Kentaro Yoshioka

机构 * Keio University(Keio大学) Tsinghua University(清华大学)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出DetAS/DetAS-X智能体框架,利用多模态大语言模型自适应组合恢复模块和专用检测器,通过自进化经验积累实现经验感知推理,在六个基准上平均F1提升28.36%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25707 2026-05-26 cs.AI 77%

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

AgentHijack:基准测试计算机使用智能体对常见环境干扰的鲁棒性

Jingwei Sun, Jianing Zhu, Yuanyi Li, Tongliang Liu, Xia HU, Bo Han

机构 * TMLR Group, Hong Kong Baptist University(香港 Baptist 大学 TMLR 团体) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Sydney AI Centre, The University of Sydney(悉尼大学 AI 中心) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 提出AgentHijack基准,通过9种可配置的常见环境干扰评估多模态大语言模型驱动的计算机使用智能体的鲁棒性,并设计AgentHijack-Agent框架提升其抗干扰能力。

Comments accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13488 2026-04-16 cs.AI 77%

Towards Scalable Lightweight GUI Agents via Multi-role Orchestration

通过多角色编排实现可扩展的轻量级GUI代理

Ziwei Wang, Junjie Zheng, Leyang Yang, Sheng Zhou, Xiaoxuan Tang, Zhouhua Fang, Zhiwei Liu, Dajun Chen, Yong Li, Jiajun Bu

机构 * Zhejiang Key Laboratory of Accessible Perception and Intelligent Systems, Zhejiang University(浙江可及感知与智能系统重点实验室,浙江大学) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) AntGroup(蚂蚁集团)

专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.AI

AI总结 本文提出LAMO框架,通过多角色编排提升轻量级GUI代理的任务扩展性,开发出支持单体执行和多代理系统编排的LAMO-3B代理,结合先进规划器实现持续性能提升。

Comments Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17043 2026-03-19 cs.CV 77%

OpenQlaw: An Agentic AI Assistant for Analysis of 2D Quantum Materials

OpenQlaw: 一种用于二维量子材料分析的代理AI助手

Sankalp Pandey, Xuan-Bac Nguyen, Hoang-Quan Nguyen, Tim Faltermeier, Nicholas Borys, Hugh Churchill, Khoa Luu

机构 * Quantum AI Lab, University of Arkansas, USA(量子人工智能实验室,美国亚拉巴马大学) University of Utah, USA(美国犹他大学) Department of Physics, University of Arkansas, USA(美国亚拉巴马大学物理系) MonARK NSF Quantum Foundry(MonARK NSF 量子熔炉)

专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CV

AI总结 OpenQlaw通过整合NanoBot和QuPAINT,实现二维材料分析中的动态推理与可视化,提升科研人员在高通量器件制造中的效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03028 2026-02-04 cs.CV 77%

MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration

MUSE:一种通过闭环认知协调的多智能体框架用于无约束故事构思

Wenzhang Sun, Zhenyu Wang, Zhangchi Hu, Chunfeng Wang, Hao Li, Wei Chen

机构 * Li Auto(利奥自动化)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 MUSE通过闭环认知协调的多智能体框架,提升长篇叙事的连贯性与多模态一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20148 2025-09-30 cs.AI 77%

MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents

Ziming Wei, Bingqian Lin, Zijian Jiao, Yunshuang Nie, Liang Ma, Yuecheng Liu, Yuzheng Zhuang, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Shanghai Jiao Tong University(上海交通大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);MLLM(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18576 2025-09-24 cs.RO cs.AI 77%

LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA

Zeyi Kang, Liang He, Yanxin Zhang, Zuheng Ming, Kaixing Zhao

机构 * School of Software Northwestern Polytechnical University Xi'an, China(软件学院 西安理工大学中国) Laboratoire L2Tl University Sorbonne Paris Nord Paris, France(L2Tl实验室 索邦巴黎北大学巴黎法国)

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02841 2025-08-06 cs.AI cs.IR 77%

A Multi-Agent System for Complex Reasoning in Radiology Visual Question Answering

Ziruo Yi, Jinyu Liu, Ting Xiao, Mark V. Albert

机构 * University of North Texas(北卡罗来纳州立大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06205 2026-07-14 cs.CL cs.AI 版本更新 76%

Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation

Tool-MCoT:用于内容安全审核的工具增强多模态链式思维

Shutong Zhang, Dylan Zhou, Yinxiao Liu, Yang Yang, Huiwen Luo, Wenfei Zou

机构 * Stanford University(斯坦福大学) Google(谷歌) Google DeepMind(谷歌DeepMind)

专题命中 多模态Agent :multimodal(title);分类 cs.CL、cs.AI

AI总结 本文提出Tool-MCoT,一种基于工具增强的多模态链式思维模型,用于提升内容安全审核的效率与准确性,通过训练小语言模型以有效利用外部工具进行推理和决策。

详情

展开后加载摘要…

URL PDF HTML 收藏