arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

2603.25195 2026-03-27 cs.HC 78%

On-Demand Instructional Material Providing Agent Based on MLLM for Tutoring Support

基于MLLM的按需教学材料提供代理用于辅导支持

Takumi Kato, Masato Kikuchi, Tadachika Ozono

专题命中 多模态Agent :MLLM(title);multimodal(abstract)

AI总结 本文提出基于多模态大语言模型的代理,用于在一对一辅导中按需提供教学材料,通过分析对话自动检索相关图片,实验显示检索时间减少44.4秒,85.7%的试次提供可接受质量的图片。

Comments The 20th International Conference on E-Service and Knowledge Management (ESKM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22447 2026-03-25 cs.SE 78%

SkillClone: Multi-Modal Clone Detection and Clone Propagation Analysis in the Agent Skill Ecosystem

SkillClone: 多模态克隆检测与克隆传播分析在智能体技能生态系统中

Jiaying Zhu, Lyuye Zhang, Wenbo Guo, Yang Liu

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出SkillClone,首个多模态智能体技能克隆检测方法,通过融合TF-IDF与通道分解实现高精度检测,揭示技能生态系统中大量克隆关系及传播问题。

Comments 12 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22507 2026-03-24 cs.NI cs.MA eess.SP 78%

A Unified Cloud-Edge-Terminal Framework for Multimodal Integrated Sensing and Communication

多模态感知与通信一体化的统一云-边-终端框架

Yubo Peng, Luping Xiang, Kun Yang, Feibo Jiang, Kezhi Wang, Christos Masouros

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出统一云-边-终端框架,通过多模态感知与通信融合,解决异构融合、通信开销和系统扩展性等挑战,提升任务导向的多模态感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01013 2026-03-24 cs.LG 78%

TimeXL: Explainable Multi-modal Time Series Prediction with LLM-in-the-Loop

TimeXL:基于LLM的多模态时间序列预测可解释方法

Yushan Jiang, Wenchao Yu, Geon Lee, Dongjin Song, Kijung Shin, Wei Cheng, Yanchi Liu, Haifeng Chen

机构 * School of Computing, University of Connecticut(大学计算机学院) Data Science & System Security Department, NEC Labs America(数据科学与系统安全部,NEC美国实验室) Kim Jaechul Graduate School of AI, KAIST(金 Jaechul人工智能研究生院,韩国科学技术院)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 TimeXL通过集成原型时间序列编码器与三个协作LLM,提升时间序列预测的准确性与可解释性,实验证明在四个真实数据集上AUC提升达8.9%。

Comments NeurIPS 2025 camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14540 2026-03-17 cs.RO 78%

Multimodal Belief-Space Covariance Steering with Active Probing and Influence for Interactive Driving

多模态信念空间协方差操控与主动探测及影响的交互驾驶

Devodita Chakravarty, John Dolan, Yiwei Lyu

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出一种多模态信念空间协方差操控方法,通过主动探测和影响策略,在交互驾驶中提升安全性和决策效率。

Comments Accepted to IEEE International Conference on Robotics and Automation (ICRA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22961 2026-03-13 stat.ME 78%

Measuring capacities in multimodal maritime port systems with anchorage queues

多模式港口系统中锚泊队列容量的测量

Debojjal Bagchi, Kyle Bathgate, Kenneth N. Mitchell, Magdalena I. Asborno, Marin M. Kress, Stephen D. Boyles

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出了一种方法,用于估算多模式港口系统的运营和终极容量,通过队列模型和微分方程模型分析休斯顿港的吞吐量及瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21100 2026-03-12 astro-ph.GA astro-ph.IM 78%

Disk Wind Feedback from High-mass Protostars. V. Application of Multi-Modal Machine Learning to Characterize Outflow Properties

高质恒星喷流反馈。V. 多模式机器学习在喷流性质表征中的应用

Duo Xu, Ioana A. Stelea, Joshua S. Speagle, Yichen Zhang, Jonathan C. Tan

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出多模式深度学习方法,通过结合空谱信息表征喷流性质,克服投影偏差,提升高质恒星形成研究的可解释性与鲁棒性。

Comments ApJ accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08336 2026-03-10 cs.RO 78%

Hierarchical Multi-Modal Planning for Fixed-Altitude Sparse Target Search and Sampling

分层多模态规划用于固定高度稀疏目标搜索与采样

Lingpeng Chen, Yuchen Zheng, Apple Pui-Yi Chui, Junfeng Wu, Ziyang Hong

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 HIMoS通过分层多模态规划提高稀疏目标搜索与采样任务的效率,结合全局与局部规划器优化路径,平衡多种传感任务。

Comments 8 pages, 9 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04363 2026-03-05 cs.RO 78%

ManipulationNet: An Infrastructure for Benchmarking Real-World Robot Manipulation with Physical Skill Challenges and Embodied Multimodal Reasoning

ManipulationNet: 一个用于基于现实世界机器人操作的基准测试基础设施,包含物理技能挑战和具身多模态推理

Yiting Chen, Kenneth Kimble, Edward H. Adelson, Tamim Asfour, Podshara Chanrungmaneekul, Sachin Chitta, Yash Chitambar, Ziyang Chen, Ken Goldberg, Danica Kragic, Hui Li, Xiang Li, Yunzhu Li, Aaron Prather, Nancy Pollard, Maximo A. Roa-Garzon, Robert Seney, Shuo Sha, Shihefeng Wang, Yu Xiang, Kaifeng Zhang, Yuke Zhu, Kaiyu Hang

机构 * Rice University(里士大学) U.S. National Institute of Standards and Technology(美国国家标准与技术研究院) Massachusetts Institute of Technology(麻省理工学院) Karlsruhe Institute of Technology(卡尔斯鲁厄技术大学) Autodesk Research(Autodesk研究) University of California, Berkeley(加州大学伯克利分校) KTH Royal Institute of Technology(皇家理工学院) Tsinghua University(清华大学) Columbia University(哥伦比亚大学) ASTM International(美国材料与试验协会) Carnegie Mellon University(卡内基梅隆大学) German Aerospace Center (DLR)(德国航空航天中心) University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Texas at Austin(德克萨斯大学奥斯汀分校) NVIDIA Research(NVIDIA研究)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 ManipulationNet通过标准化硬件和统一软件客户端,为机器人操作提供现实世界基准测试,促进物理技能和具身推理能力的系统性发展。

Comments 32 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02635 2026-03-04 cs.LG 78%

SaFeR-ToolKit: Structured Reasoning via Virtual Tool Calling for Multimodal Safety

SaFeR-ToolKit: 通过虚拟工具调用实现多模态安全的结构化推理

Zixuan Xu, Tiancheng He, Huahui Yi, Kun Wang, Xi Chen, Gongli Xi, Qiankun Li, Kang Li, Yang Liu, Zhigang Zeng

机构 * Huazhong University of Science and Technology(华中科技大学) Beijing University of Posts and Telecommunications(北京邮电大学) West China Hospital, Sichuan University(四川大学华西医院) Nanyang Technological University(南洋理工大学)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 SaFeR-ToolKit通过虚拟工具调用实现多模态安全的结构化推理,提升安全性、帮助性和推理严谨性,同时保持通用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02503 2026-03-04 eess.SY cs.SY 78%

Joint Estimation of Dynamic O-D Demand and Choice Models for Dynamic Multi-modal Networks: Computational Graph-Based Learning and Hypothesis Tests

动态多模式网络中动态O-D需求与选择模型的联合估计:基于计算图的学习与假设检验

Xiaoyu Ma, Sean Qian

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本研究提出基于计算图的学习方法,联合估计多模式网络中动态O-D需求与选择模型,通过假设检验框架提升模型的统计显著性分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21157 2026-03-02 cs.RO 78%

HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning

HALO:一种用于具身多模态推理的统一视觉-语言-动作模型

Quanxin Shou, Fangqi Zhu, Shawn Chen, Puxin Yan, Zhengyang Yan, Yikun Miao, Xiaoyi Pang, Zicong Hong, Ruikai Shi, Hao Huang, Jie Zhang, Song Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 HALO提出了一种统一的视觉-语言-动作模型,通过结合文本推理、视觉子目标预测和增强的动作预测,实现具身多模态链式推理,提升了机器人在复杂环境中的表现和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19346 2026-02-24 cs.RO cs.SY eess.SY 78%

Design and Control of Modular Magnetic Millirobots for Multimodal Locomotion and Shape Reconfiguration

模块化磁性微机器人多模态运动与形状重构设计与控制

Erik Garcia Oyono, Jialin Lin, Dandan Zhang

机构 * Department of Bioengineering, Imperial College London(帝国理工学院生物工程系)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本研究提出一种模块化磁性微机器人平台,通过多模块协同实现多模态运动与形状重构,展示了在受限环境中稳健控制的潜力。

Comments Accepted by 2026 ICRA

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08417 2026-02-17 cs.RO 78%

ORACLE-Grasp: Zero-Shot Affordance-Aligned Robotic Grasping using Large Multimodal Models

ORACLE-Grasp: 基于大多模态模型的零样本 affordance 对齐机器人抓取

Avihai Giuili, Rotem Atari, Avishai Sintov

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 ORACLE-Grasp 利用大多模态模型实现零样本抓取,通过语义对齐提升抓取准确性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11057 2026-02-12 cs.LG 78%

Divide, Harmonize, Then Conquer It: Shooting Multi-Commodity Flow Problems with Multimodal Language Models

分割、调和、然后征服:利用多模态语言模型解决多商品流问题

Xinyu Yuan, Yan Qiao, Zonghui Wang, Wenzhi Chen

机构 * Zhejiang University(浙江大学) Hefei University of Technology(合肥工业大学) Co-corresponding authors(共同通讯作者)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 Pram利用多模态语言模型解决多商品流问题,通过分解和调和子问题实现高效优化,性能接近线性规划求解器且运行时间显著降低。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05671 2026-02-06 cs.HC 78%

(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences

行动中的视觉:比较远程视觉援助与多模态语音代理在检查序列中的表现

Damien Rudaz, Barbara Nino Carreras, Sara Merlino, Brian L. Due, Barry Brown

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 研究比较了远程视觉援助与多模态语音代理在检查任务中的表现,发现代理无法产生基于环境的视觉动作,从而缺乏关键资源。

Comments Conditionally accepted at CHI 2026, 32 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04157 2026-02-05 cs.RO 78%

A Modern System Recipe for Situated Embodied Human-Robot Conversation with Real-Time Multimodal LLMs and Tool-Calling

一种面向情境具身人机对话的现代系统配方,结合实时多模态大语言模型与工具调用

Dong Won Lee, Sarah Gillet, Louis-Philippe Morency, Cynthia Breazeal, Hae Won Park

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出了一种结合实时多模态大语言模型与工具调用的系统配方,用于提升情境具身人机对话的交互质量与效率。

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19185 2026-02-03 math.OC 78%

Distributionally Robust Optimization with Multimodal Decision-Dependent Ambiguity Sets

具有多模决策依赖不确定集的分布鲁棒优化

Xian Yu, Beste Basciftci

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出了一种基于ϕ-分歧度的多模决策依赖分布鲁棒优化框架,通过引入多模不确定集和分解算法,改进了DRO模型的求解方法,并通过案例验证了多模性对优化性能的提升作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03777 2026-01-08 stat.ME cs.SY eess.SY 78%

Multi-agent Optimization of Non-cooperative Multimodal Mobility Systems

非合作多模式移动系统的多智能体优化

Md Nafees Fuad Rafi, Zhaomiao Guo

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 本文提出了一种多智能体优化框架,用于分析非合作多模式移动系统中旅行者和司机的市场互动,通过均衡定价平衡供需并优化系统效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00696 2026-01-05 cs.LG cs.GT cs.RO 78%

Bayesian Inverse Games with High-Dimensional Multi-Modal Observations

高维多模态观测下的贝叶斯逆游戏

Yash Jain, Xinjie Liu, Lasse Peters, David Fridovich-Keil, Ufuk Topcu

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Delft University of Technology(代尔夫特理工大学) Sunrise Setting Ltd SAGE Publications Ltd(SAGE出版社有限公司)

专题命中 多模态Agent :multi-modal(title);multimodal(abstract)

AI总结 本文提出了一种基于贝叶斯推断的逆博弈框架,利用多模态观测数据实时生成隐藏智能体目标的后验分布,提升推断质量并实现更安全的决策。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06115 2026-01-05 cs.RO cs.SY eess.SY 78%

Hybrid A* Path Planning with Multi-Modal Motion Extension for Four-Wheel Steering Mobile Robots

四轮转向移动机器人多模态运动扩展的混合A*路径规划

Runjiao Bao, Lin Zhang, Tianwei Niu, Haoyu Yuan, Shoukun Wang

机构 * School of Automation, Beijing Institute of Technology, Beijing 100081, China(自动化学院,北京理工大学,北京100081,中国)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出了一种针对四轮转向移动机器人的混合A*路径规划方法,通过多模态运动扩展提升复杂环境下的路径规划性能。

Comments Updated method details, parameters, and experimental scenarios

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19742 2025-12-24 cs.LG 78%

On-device Large Multi-modal Agent for Human Activity Recognition

用于人体活动识别的设备端大多模态智能体

Md Shakhrul Iman Siam, Ishtiaque Ahmed Showmik, Guanqun Song, Ting Zhu

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出了一种用于人体活动识别的设备端大多模态智能体,结合大语言模型提升性能与可解释性,实现高分类准确率和用户友好交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11876 2025-12-16 cs.RO cs.SY eess.SY 78%

Traversability Aware Autonomous Navigation for Multi-Modal Mobility Morphobot (M4)

多模态移动形变机器人(M4)的可 traversability 自主导航

Hrigved Mahesh Suryawanshi

机构 * SiliconSynapse Lab(硅合成实验室)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本研究提出了一种基于LiDAR的多模态移动形变机器人M4的可 traversability 自主导航框架,通过学习地形分析生成节能路径,提升地形适应能力。

Comments Master's thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07132 2025-12-09 cs.CL cs.AI cs.CV 78%

DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning

利用多智能体分歧进行多模态推理中的工具招募

Nithin Sivakumaran, Justin Chih-Yao Chen, David Wan, Yue Zhang, Jaehong Yoon, Elias Stengel-Eskin, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 多模态Agent :multimodal(title);分类 cs.CV、cs.CL、cs.AI

AI总结 DART通过多智能体分歧识别有用视觉工具,提升多模态推理中的工具调用效果。

Comments Code: https://github.com/nsivaku/dart

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04308 2025-12-05 cs.RO 78%

ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models

ResponsibleRobotBench: 使用多模态大语言模型评估负责任的机器人操控

Lei Zhang, Ju Dong, Kaixin Bai, Minheng Ni, Zoltan-Csaba Marton, Zhaopeng Chen, Jianwei Zhang

机构 * University of Hamburg(汉堡大学) Agile Robots SE(敏捷机器人公司) Technical University of Munich(慕尼黑技术大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态Agent :multi-modal(title);multimodal(abstract)

AI总结 ResponsibleRobotBench通过多模态大语言模型评估机器人操控的责任性,涵盖23个多阶段任务,强调安全性、泛化能力和物理可靠性。

Comments https://sites.google.com/view/responsible-robotbench

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02777 2025-12-03 cs.RO cs.MA 78%

CogDrive: Cognition-Driven Multimodal Prediction-Planning Fusion for Safe Autonomy

CogDrive: 基于认知的多模态预测-规划融合用于安全自主性

Heye Huang, Yibin Yang, Mingfeng Fan, Haoran Wang, Xiaocong Zhao, Jianqiang Wang

机构 * Singapore-MIT Alliance for Research and Technology (SMART), Singapore(新加坡-麻省理工联合研究技术联盟) Department of Urban Studies and Planning, Massachusetts Institute of Technology, USA(麻省理工学院城市研究与规划系) School of Vehicle and Mobility, Tsinghua University, China(清华大学车辆与移动系统学院) Department of Mechanical Engineering, National University of Singapore, Singapore(新加坡国立大学机械工程系) Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University, China(同济大学交通工程重点实验室)

专题命中 多模态Agent :multimodal(title,abstract)

AI总结 CogDrive通过结合认知多模态预测与安全导向规划,实现了在复杂交通中的安全自主性,提升了轨迹预测和适应性行为。

Comments 25 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01310 2025-11-12 cs.MA 78%

From Pixels to Cooperation Multi Agent Reinforcement Learning based on Multimodal World Models

Sureyya Akin, Kavita Srivastava, Prateek B. Kapoor, Pradeep G. Sethi, Sunita Q. Patel, Rahu Srivastava

专题命中 多模态Agent :multimodal(title,abstract)

Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18515 2025-11-12 cs.MA 78%

Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models

Sureyya Akin, Shruti T. Tiwari, Ram Bhattacharya, Sagar A. Raman, Kiran Mohanty, Sita Krishnan

专题命中 多模态Agent :multimodal(title,abstract)

Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07315 2025-11-11 cs.CR 78%

JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework

Yuxuan Zhou, Yang Bai, Kuofeng Gao, Tao Dai, Shu-Tao Xia

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07219 2025-11-11 q-bio.GN 78%

Integrating Epigenetic and Phenotypic Features for Biological Age Estimation in Cancer Patients via Multimodal Learning

Shuyue Jiang, Wenjing Ma, Shaojun Yu, Chang Su, Runze Yan, Jiaying Lu

专题命中 多模态Agent :multimodal(title,abstract)

Journal ref In Proceedings of The 19th IEEE International Conference on Bioinformatics and Biomedicine (BIBM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏