arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-10 至 2026-03-10 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 14 篇

2510.17277 2026-03-10 cs.CR 88%

PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs

PolyJailbreak: 对黑盒多模态大语言模型的跨模态劫持攻击

Xinkai Wang, Beibei Li, Zerui Shao, Ao Liu, Guangquan Xu, Shouling Ji

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(title,abstract)

AI总结 PolyJailbreak通过多模态安全性不对称现象,提出结构化原子策略原语库,实现对黑盒多模态大语言模型的高效劫持攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08240 2026-03-10 cs.CV 83%

SiMO: Single-Modality-Operable Multimodal Collaborative Perception

SiMO: 单模态可操作的多模态协作感知

Jiageng Wen, Shengjie Zhao, Bing Li, Jiafeng Huang, Kenan Ye, Hao Deng

机构 * Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学智能自主系统上海研究院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) School of Mechatronic Engineering and Automation, Shanghai University(上海大学机械电子工程与自动化学院)

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 SiMO通过单模态可操作的多模态协作感知方法,解决多模态特征融合中的语义不匹配问题,提升协作感知性能。

Comments Accepted to ICLR 2026. This arXiv version includes an additional appendix (Appendix 15) containing further philosophical discussion not included in the official ICLR peer-reviewed version

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08369 2026-03-10 cs.AI 79%

M$^3$-ACE: Rectifying Visual Perception in Multimodal Math Reasoning via Multi-Agentic Context Engineering

M$^3$-ACE: 通过多智能体上下文工程校正多模态数学推理中的视觉感知

Peijin Xie, Zhen Xu, Bingquan Liu, Baoxun Wang

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

AI总结 M3-ACE 通过多智能体上下文工程校正多模态数学推理中的视觉感知问题,提升推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08336 2026-03-10 cs.RO 78%

Hierarchical Multi-Modal Planning for Fixed-Altitude Sparse Target Search and Sampling

分层多模态规划用于固定高度稀疏目标搜索与采样

Lingpeng Chen, Yuchen Zheng, Apple Pui-Yi Chui, Junfeng Wu, Ziyang Hong

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 HIMoS通过分层多模态规划提高稀疏目标搜索与采样任务的效率,结合全局与局部规划器优化路径,平衡多种传感任务。

Comments 8 pages, 9 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08260 2026-03-10 cs.RO 71%

Seed2Scale: A Self-Evolving Data Engine for Embodied AI via Small to Large Model Synergy and Multimodal Evaluation

Seed2Scale: 一种通过小到大模型协同和多模态评估的自我进化数据引擎用于具身AI

Cong Tai, Zhaoyu Zheng, Haixu Long, Hansheng Wu, Zhengbin Long, Haodong Xiang, Rong Shi, Zhuo Cui, Shizhuang Zhang, Gang Qiu, He Wang, Ruifeng Li, Biao Liu, Zhenzhe Sun, Tao Shen

机构 * ZTE Corporation(中兴通讯公司)

专题命中 多模态Agent :multimodal(title)

AI总结 Seed2Scale通过小到大模型协同和多模态评估,实现具身AI的自我进化数据引擎,显著提升性能并提供可扩展的开发路径

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18064 2026-03-10 cs.CV 70%

3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis

3DMedAgent:面向3D医学分析的统一感知-理解框架

Ziyue Wang, Linghan Cai, Chang Han Low, Haofeng Liu, Junde Wu, Jingyu Wang, Rui Wang, Lei Song, Jiang Bian, Jingjing Fu, Yueming Jin

机构 * National University of Singapore(新加坡国立大学) Microsoft Research(微软研究院) TUD Dresden University of Technology(德累斯顿技术大学) University of Oxford(牛津大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 3DMedAgent通过统一框架实现2D MLLMs在3D医学分析中的高效处理,无需3D特定微调,提升3D医学图像的感知与理解能力。

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02951 2026-03-10 cs.LG cs.CV 70%

CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning

CGL: 通过强化微调推进连续GUI学习

Zhenquan Yao, Zitong Huang, Yihan Zeng, Jianhua Han, Hang Xu, Chun-Mei Feng, Jianwei Ma, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨工业大学) Huawei Noah’s Ark Lab(华为诺亚实验室) University College Dublin(都柏林大学) Peking University(北京大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 CGL通过强化微调与监督微调的协同优化,解决GUI连续学习中遗忘旧任务的问题,提出AndroidControl-CL基准验证方法有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14899 2026-03-10 cs.AI cs.CV 62%

InsightX Agent: An LMM-based Agentic Framework with Integrated Tools for Reliable X-ray NDT Analysis

InsightX Agent: 一种基于LMM的智能框架,集成工具用于可靠X射线无损检测分析

Jiale Liu, Huan Wang, Yue Zhang, Xiaoyu Luo, Jiaxiang Hu, Zhiliang Liu, Min Xie

机构 * School of Physics and Astronomy, The University of Edinburgh(物理学与天文学学院,爱丁堡大学) Glasgow College, University of Electronic Science and Technology of China(电子科技大学成都学院) School of Mechanical and Electrical Engineering, University of Electronic Science and Technology of China(机械与电子工程学院,电子科技大学) Department of Systems Engineering, City University of Hong Kong(系统工程系,香港城市大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 InsightX Agent基于LMM构建智能框架,集成工具提升X射线检测的可靠性与可解释性,实现主动推理与高质量分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02083 2026-03-10 cs.RO cs.CV 57%

$π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs

$π$-StepNFT: 更宽的空间需要更细的步骤在线RL用于基于流的VLAs

Siting Wang, Xiaofeng Wang, Zheng Zhu, Minnan Pei, Xinyu Cui, Cheng Deng, Jian Zhao, Guan Huang, Haifeng Zhang, Jun Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 $π$-StepNFT通过分步负向感知微调方法,在在线强化学习中提升基于流的VLAs的性能和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08113 2026-03-10 cs.CV 57%

SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving

SAMoE-VLA:一种面向自动驾驶的场景自适应混合专家视觉-语言-动作模型

Zihan You, Hongwei Liu, Chenxu Dang, Zhe Wang, Sining Ang, Aoqi Wang, Yan Wang

机构 * Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学) School of Instrument Science and Engineering, Southeast University(仪器科学与工程学院,东南大学) Zhili College, Tsinghua University(紫荆学院,清华大学) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) Department of Automation, University of Science and Technology of China(自动化学院,中国科学技术大学) Department of Automation, University of Science and Technology Beijing(自动化学院,北京科技大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

AI总结 SAMoE-VLA通过场景自适应混合专家机制提升自动驾驶中的视觉-语言-动作推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08013 2026-03-10 cs.AI 57%

PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation Agents

PIRA-Bench: 从反应式 GUI 代理到基于 GUI 的主动意图推荐代理的转变

Yuxiang Chai, Shunye Tang, Han Xiao, Rui Liu, Hongsheng Li

机构 * Nankai University(南开大学) Huawei Research(华为研究)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 PIRA-Bench通过引入主动意图推荐代理基准,推动GUI代理从反应式向主动式转变,评估多模态大语言模型在连续弱监督视觉输入中的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10918 2026-03-10 cs.HC cs.CL 57%

CompanionCast: Toward Social Collaboration with Multi-Agent Systems in Shared Experiences

CompanionCast:迈向多智能体系统在共享体验中的社会协作

Yiyang Wang, Chen Chen, Tica Lin, Vishnu Raj, Josh Kimball, Alex Cabral, Josiah Hester

机构 * Georgia Institute of Technology(佐治亚理工学院) Dolby Laboratories, Inc.(杜比实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

AI总结 CompanionCast通过多智能体系统提升共享体验中的社会协作,通过多模态检测、上下文缓存和空间音频增强共在感,实验显示其在体育观看中显著提升社会存在感和情感共享。

Comments Accepted at ACM CHI 2026 Workshop on Human-Agent Collaboration

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11978 2026-03-10 cs.RO cs.AI 57%

Accelerating Robotic Reinforcement Learning with Agent Guidance

通过代理引导加速机器人强化学习

Haojun Chen, Zili Zou, Chengdong Ma, Yaoxiang Pu, Haotong Zhang, Yuanpei Chen, Yaodong Yang

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) PKU-PsiBot Joint Lab(北京大学- PsiBot 联合实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 AGPS通过多模态代理替代人类监督,提升机器人强化学习的样本效率和可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09587 2026-03-10 cs.RO 50%

GeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation

GeoNav: 通过双尺度地理推理赋能大语言模型实现语言目标的空中导航

Haotian Xu, Yue Hu, Chen Gao, Zhengqiu Zhu, Yong Zhao, Yong Li, Quanjun Yin

机构 * College of Systems Engineering, National University of Defense Technology(系统工程学院,国防科技大学) State Key Laboratory of Digital Intelligent Modeling and Simulation(数字智能建模与仿真国家重点实验室) BNRist, Tsinghua University(清华大学BNRist)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 GeoNav通过双尺度地理推理赋能大语言模型,实现语言目标的空中导航,提升城市场景下的导航成功率和精度。

Comments Published in Pattern Recognition (2026)

Journal ref Pattern Recognition, Volume 177, 113365, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏