arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2783 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2783 篇

2601.02377 2026-01-07 cs.RO 50%

Trust in LLM-controlled Robotics: a Survey of Security Threats, Defenses and Challenges

对LLM控制的机器人信任:安全威胁、防御和挑战的综述

Xinyu Huang, Shyam Karthick V B, Taozhao Chen, Mitch Bryson, Thomas Chaffey, Huaming Chen, Kim-Kwang Raymond Choo, Ian R. Manchester

机构 * School of Electrical and Computer Engineering, The University of Sydney(悉尼大学电气与计算机工程学院) Australian Centre for Robotics and School of Aerospace, Mechanical and Mechatronic Engineering, The University of Sydney(悉尼大学机器人中心及航空航天、机械与机电工程学院) Department of Information Systems and Cybersecurity, University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校信息系统与网络安全系) School of Engineering and Natural Sciences, University of Iceland(冰岛大学工程与自然科学学院)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文综述了LLM控制机器人面临的安全威胁和防御策略,强调了具身系统中恶意输出的物理风险,并提出了安全规范和多LLM监督等防御方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16760 2026-01-06 cs.RO 50%

Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future

面向自动驾驶的视觉-语言-动作模型:过去、现在与未来

Tianshuai Hu, Xiaolu Liu, Song Wang, Yiyao Zhu, Ao Liang, Lingdong Kong, Guoyang Zhao, Zeying Gong, Jun Cen, Zhiyu Huang, Xiaoshuai Hao, Linfeng Li, Hang Song, Xiangtai Li, Jun Ma, Shaojie Shen, Jianke Zhu, Dacheng Tao, Ziwei Liu, Junwei Liang

机构 * HKUST(香港科技大学) Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) HKUST(GZ)(香港科技大学(广州)) DAMO Academy, Alibaba(阿里巴巴达摩院) University of California, Los Angeles(加州大学洛杉矶分校) Xiaomi EV(小米电动车) Xi'an Jiaotong University(西安交通大学) Nanyang Technological University, Singapore(新加坡南洋理工大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文探讨了自动驾驶中视觉-语言-动作模型的发展历程,提出两种主要范式并分析其挑战与未来方向。

Comments Survey; 47 pages, 7 figures, 9 tables; GitHub Repo at https://github.com/worldbench/awesome-vla-for-ad

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24645 2026-01-01 cs.SD 50%

AudioFab: Building A General and Intelligent Audio Factory through Tool Learning

通过工具学习构建通用且智能的音频工厂

Cheng Zhu, Jing Han, Qianshuai Xue, Kehan Wang, Huan Zhao, Zixing Zhang

机构 * Hunan University(湖南大学) University of Cambridge(剑桥大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 AudioFab通过工具学习构建通用智能音频处理框架,提供模块化设计和自然语言接口,提升音频任务的效率和准确性。

Journal ref ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23158 2025-12-30 eess.SY cs.RO cs.SY 50%

Breaking Symmetry-Induced Degeneracy in Multi-Agent Ergodic Coverage via Stochastic Spectral Control

通过随机频谱控制打破多智能体等耗覆盖中的对称诱导退化

Kooktae Lee, Julian Martinez

机构 * Department of Mechanical Engineering, New Mexico Institute of Mining and Technology(机械工程系,新墨西哥矿业与技术研究所)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出随机频谱控制方法,通过引入随机扰动和收缩项,解决多智能体等耗覆盖中因对称性导致的梯度抵消问题,确保轨迹有界并避免停滞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22968 2025-12-30 eess.SY cs.SY 50%

A Bezier Curve Based Approach to the Convexification of the AC Optimal Power Flow Problem

基于贝塞尔曲线的AC最优潮流问题凸化方法

Carlos Arturo Saldarriaga-Cortes, Carlos Adrian Correa-Florez, Maximiliano Bueno-Lopez, Maria Victoria Gasca-Segura

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出基于贝塞尔曲线的AC最优潮流问题凸化方法,通过引入辅助变量和对数变换,实现高效且准确的电力系统优化。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19854 2025-12-30 cs.RO cs.HC 50%

Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models

思考、行动、学习:一种利用闭环大语言模型的自主机器人代理框架

Anjali R. Menon, Rohit K. Sharma, Priya Singh, Chengyu Wang, Aurora M. Ferreira, Mateja Novak

机构 * Dept. of Electronics & Comm. Eng.(电子与通信工程系) Government Engineering College(政府工程学院) Dept. of Electrical & Electronics Eng.(电气与电子工程系) Poornima College of Engineering(波奥尼玛工程学院) Dept. of Electronics & Telecom. Eng.(电子与电信工程系) Shivaji University College of Eng.(希瓦吉大学工程学院) Department of Computer Science(计算机科学系) San Francisco State University(旧金山州立大学) Dept. of Electrical Eng.(电气工程系) Instituto Federal do Maranhão(马里兰联邦学院) Dept. of Electrical & Computer Eng.(电气与计算机工程系) Technical University of Košice(科希丘夫技术大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出T-A-L框架,通过闭环大语言模型实现机器人自主学习与策略优化,显著提升复杂任务的成功率和泛化能力。

Comments 13 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13537 2025-12-23 math.LO 50%

Models for the common knowledge logic

共同知识逻辑的模型

Yoshihito Tanaka

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文研究了共同知识逻辑的模型,探讨了其框架和代数的定义及性质,指出CKL框架是模态可定义的,而CKL代数并非构成变种。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11620 2025-12-15 cs.RO cs.SY eess.SY 50%

Architecting Large Action Models for Human-in-the-Loop Intelligent Robots

为具有人类在环的智能机器人构建大型动作模型

Kanisorn Sangchai, Methasit Boonpun, Withawin Kraipetchara, Paulo Garcia

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出通过整合符号方法与现成模型构建可验证的神经符号智能机器人动作模型,以提升可靠性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08476 2025-12-10 cs.RO 50%

A Multi-Agent LLM Framework for Design Space Exploration in Autonomous Driving Systems

多智能体大语言模型框架用于自动驾驶系统的设计空间探索

Po-An Shih, Shao-Hua Wang, Yung-Che Li, Chia-Heng Tu, Chih-Han Chang

机构 * National Cheng Kung University(国立成功大学) Safeware Technology Inc.(Safeware技术公司)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出基于多智能体大语言模型的自动驾驶系统设计空间探索框架,通过多模态推理和3D模拟工具自动化解析执行输出,提升设计效率和成本效益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08186 2025-12-10 cs.RO 50%

Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation

放慢脚步,快速移动:一种通用视觉-语言导航的双系统基础模型

Meng Wei, Chenyang Wan, Jiaqi Peng, Xiqian Yu, Yuqiang Yang, Delin Feng, Wenzhe Cai, Chenming Zhu, Tai Wang, Jiangmiao Pang, Xihui Liu

机构 * Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Zhejiang University(浙江大学) Tsinghua University(清华大学)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 DualVLN通过双系统架构,结合高层推理与低层动作执行,提升视觉-语言导航的泛化能力和动态环境适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23823 2025-12-10 cs.RO 50%

Control Your Robot: A Unified System for Robot Control and Policy Deployment

掌控你的机器人:一种统一的机器人控制与策略部署系统

Tian Nian, Weijie Ke, Shaolong Zhu, Bingshan Hu

机构 * ScaleLab, Shanghai Jiao Tong University(上海交通大学ScaleLab) University of Shanghai for Science and Technology(上海科学技术大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 Control Your Robot提出了一种统一的机器人控制与策略部署框架,通过模块化设计和标准化流程,实现跨平台的数据收集与策略学习,提升机器人学习的可扩展性和可重复性。

Comments Code: https://github.com/Tian-Nian/control_your_robot

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07452 2025-12-09 cs.IR 50%

From Show Programmes to Data: Designing a Workflow to Make Performing Arts Ephemera Accessible Through Language Models

从节目到数据:设计一种工作流,通过语言模型使表演艺术的临时性资料可访问

Clarisse Bardiot, Pierre-Carl Langlais, Bernard Jacquemin, Jacob Hart, Antonios Lagarias, Nicolas Foucault, Aurélie Lemaître-Legargeant, Jeanne Fras

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出利用多模态语言模型和本体推理模型,将戏剧节目转化为结构化数据,以提升文化遗产资料的可访问性和分析能力。

Comments 19 pages, 8 figures, 5 tables, 17 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07032 2025-12-09 cs.RO 50%

A Hetero-Associative Sequential Memory Model Utilizing Neuromorphic Signals: Validated on a Mobile Manipulator

一种利用神经形态信号的异关联序列记忆模型:在移动机械臂上的验证

Runcong Wang, Fengyi Wang, Gordon Cheng

机构 * Institute for Cognitive Systems, Technical University of Munich(认知系统研究所,慕尼黑技术大学)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出一种利用神经形态信号的异关联序列记忆模型,用于移动机械臂的触觉与动作决策,实现低计算成本的伪柔顺控制与多关节抓取序列检索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06558 2025-12-09 cs.RO 50%

Embodied Referring Expression Comprehension in Human-Robot Interaction

具身指称表达理解在人机交互中的应用

Md Mofijul Islam, Alexi Gladstone, Sujan Sarker, Ganesh Nanduru, Md Fahim, Keyan Du, Aman Chadha, Tariq Iqbal

机构 * University of Virginia(弗吉尼亚大学) Stanford University(斯坦福大学) University of Dhaka(达卡大学) Amazon GenAI(亚马逊生成人工智能)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出Refer360数据集和MuRes模块,用于提升机器人在人机交互中对具身指称表达的理解能力。

Comments 14 pages, 7 figures, accepted at the ACM/IEEE International Conference on Human-Robot Interaction (HRI) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22487 2025-12-08 eess.SY cs.SY 50%

A Multi-Objective Simultaneous Routing, Facility Location and Allocation Model for Earthquake Emergency Logistics

面向地震应急物流的多目标同时路由、设施选址与分配模型

Sakineh Khodadadi, Tohid Kargar Tasooji, Afshin Shariat-Mohayman, Navid Kalantari

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出多目标优化模型,同时优化地震应急物流中的路由、设施选址与医院分配,以减少需求未满足、伤员未服务和经济成本。

Comments One of the authors does not agree to publish the paper on arXiv

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22390 2025-12-01 cs.LO 50%

Modal Logic for Simulation, Refinement, and Mutual Ignorance

模态逻辑用于模拟、细化和相互无知

Hans van Ditmarsch, Tim French, Rustam Galimullin, Louwe B. Kuijer

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出了一种基于多模态逻辑的模态逻辑,用于模拟、细化和相互无知,通过模块化的方式构建了多种逻辑体系,探讨了细化与模拟之间的关系。

Comments In Proceedings TARK 2025, arXiv:2511.20540

Journal ref EPTCS 437, 2025, pp. 379-398

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21000 2025-11-27 cs.HC 50%

PileUp: A Tufting Approach to Soft, Tactile, and Volumetric E-Textile Interfaces

PileUp:一种基于毛毡的柔软、触觉和体积电子织物接口方法

Seoyoung Choi, Rashmi Balegar Mohan, Heather Jin Hee Kim, Jisoo Ha, Jeyeon Jo

专题命中 多模态Agent :multimodal(abstract)

AI总结 PileUp通过毛毡技术实现柔软、触觉和体积电子织物接口,能够检测压力、弯曲、应变及湿度等,为集成传感织物提供新的制造方法。

Comments Twentieth International Conference on Tangible, Embedded, and Embodied Interaction (TEI '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20429 2025-11-26 astro-ph.CO astro-ph.GA 50%

Estimating the triaxiality of massive clusters from 2D observables in MillenniumTNG with machine learning

从MillenniumTNG模拟中的2D观测数据利用机器学习估计巨团的三轴性

Ana Maria Delgado, Michelle Ntampaka, Sownak Bose, Fulvio Ferlito, Boryana Hadzhiyska, Lars Hernquist, John Soltis, John F. Wu, Mikaeel Yunus, John ZuHone

专题命中 多模态Agent :multi-modal(abstract)

AI总结 利用机器学习从MillenniumTNG模拟中的2D观测数据估计巨团的三轴性和取向,提升团几何估计精度30%

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.12963 2025-11-26 q-bio.NC q-bio.QM 50%

Modeling the subjective perspective of consciousness and its role in the control of behaviours

意识的主观视角建模及其在行为控制中的作用

D. Rudrauf, G. Sergeant-Perthuis, O. Belli, Y. Tisserand, G. Di Marzo Serugendo

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出基于投影意识模型的主观视角建模,揭示其在行为控制中的作用,通过智能体模拟展示适应与非适应性行为的生成。

Journal ref Journal of Theoretical Biology Volume 534, 7 February 2022, 110957

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16063 2025-11-21 cs.NI 50%

Modeling Pointing, Acquisition, and Tracking Delays in Free-Space Optical Satellite Networks

自由空间光学卫星网络中指针、获取和跟踪延迟的建模

Jason Gerard, Juan A. Fraire, Sandra Céspedes

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种用于自由空间光学卫星网络中指针、获取和跟踪延迟建模的验证模型,以提高接触计划的准确性和网络利用率。

Comments 2025 IEEE International Conference on Wireless for Space and Extreme Environments (WiSEE) - STINT Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11120 2025-11-18 cs.CR 50%

SoK: How Sensor Attacks Disrupt Autonomous Vehicles: An End-to-end Analysis, Challenges, and Missed Threats

Qingzhao Zhang, Shaocheng Luo, Z. Morley Mao, Miroslav Pajic, Michael K. Reiter

专题命中 多模态Agent :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19838 2025-11-18 cs.HC 50%

LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects

Guangyi Liu, Pengxiang Zhao, Yaozhen Liang, Liang Liu, Yaxuan Guo, Han Xiao, Weifeng Lin, Yuxiang Chai, Yue Han, Shuai Ren, Hao Wang, Xiaoyu Liang, WenHao Wang, Tianze Wu, Zhengxi Lu, Siheng Chen, LiLinghao, Hao Wang, Guanjing Xiong, Yong Liu, Hongsheng Li

专题命中 多模态Agent :multimodal(abstract)

Comments Paper accepted to TMLR 2025, Project Homepage: https://github.com/PhoneLLM/Awesome-LLM-Powered-Phone-GUI-Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23693 2025-11-13 physics.flu-dyn 50%

CFDagent: A Language-Guided, Zero-Shot Multi-Agent System for Complex Flow Simulation

Zhaoyue Xu, Long Wang, Chunyu Wang, Yixin Chen, Qingyong Luo, Hua-Dong Yao, Shizhao Wang, Guowei He

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12844 2025-11-13 cs.IR 50%

Machine-Readable Ads: Accessibility and Trust Patterns for AI Web Agents interacting with Online Advertisements

Joel Nitu, Heidrun Mühle, Andreas Stöckl

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08661 2025-11-13 cs.RO cs.SY eess.SY 50%

SafeFlow: Safe Robot Motion Planning with Flow Matching via Control Barrier Functions

Xiaobing Dai, Zewen Yang, Dian Yu, Fangzhou Liu, Hamid Sadeghian, Sami Haddadin, Sandra Hirche

机构 * Chair of Information-oriented Control (ITR), School of Computation, Information and Technology (CIT), Technical University of Munich (TUM)(信息导向控制教授职位、计算信息与技术学院(CIT)、慕尼黑技术大学(TUM)) Chair of Robotics and Systems Intelligence, Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM)(机器人与系统智能教授职位、慕尼黑机器人与机器智能研究所(MIRMI)、慕尼黑技术大学(TUM)) Professorship of AI Planning in Dynamic Environments, School of Computation, Information and Technology (CIT), Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM)(动态环境中人工智能规划教授职位、计算信息与技术学院(CIT)、慕尼黑机器人与机器智能研究所(MIRMI)、慕尼黑技术大学(TUM)) National Key Laboratory of Modeling and Simulation for Complex Systems, Harbin Institute of Technology(复杂系统建模与仿真国家重点实验室、哈尔滨工业大学)

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24030 2025-11-12 cs.MA 50%

Human Machine Social Hybrid Intelligence:A Collaborative Decision Making Framework for Large Model Agent Groups and Human Experts

Ahmet Akkaya Melih, Yamuna Singh, Kunal L. Agarwal, Priya Mukherjee, Kiran Pattnaik, Hanuman Bhatia

专题命中 多模态Agent :multi-modal(abstract)

Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01386 2025-11-12 cs.LG cs.AR 50%

CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization

Irene Wang, Newsha Ardalani, Mostafa Elhoushi, Daniel Jiang, Samuel Hsia, Ekin Sumbul, Divya Mahajan, Carole-Jean Wu, Bilge Acun

机构 * Georgia Institute of Technology(佐治亚理工学院) FAIR at Meta(Meta的FAIR部门) Reality Labs at Meta(Meta的Reality Labs) Meta

专题命中 多模态Agent :multi-modal(abstract)

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05912 2025-11-11 eess.SP 50%

RadioSim Agent: Combining Large Language Models and Deterministic EM Simulators for Interactive Radio Map Analysis

Sajjad Hussain, Conor Brennan

专题命中 多模态Agent :multimodal(abstract)

Comments Submitted to EuCAP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05723 2025-11-11 cs.RO 50%

TumorMap: A Laser-based Surgical Platform for 3D Tumor Mapping and Fully-Automated Tumor Resection

Guangshen Ma, Ravi Prakash, Beatrice Schleupner, Jeffrey Everitt, Arpit Mishra, Junqin Chen, Brian Mann, Boyuan Chen, Leila Bridgeman, Pei Zhong, Mark Draelos, William C. Eward, Patrick J. Codd

机构 * Thomas Lord Department of Mechanical Engineering and Materials Science, Duke University(杜克大学机械工程与材料科学系) Department of Robotics, University of Michigan, Ann Arbor(密歇根大学机器人学系) Department of Orthopaedic Surgery, School of Medicine, Duke University(杜克大学骨科手术系) Department of Pathology, School of Medicine, Duke University(杜克大学病理学系) Department of Ophthalmology and Visual Sciences, University of Michigan Medical School, Ann Arbor(密歇根大学医学学院眼科与视觉科学系) Department of Neurosurgery, School of Medicine, Duke University(杜克大学神经外科系)

专题命中 多模态Agent :multimodal(abstract)

Comments 41 pages, 25 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15922 2025-11-11 cs.LG cs.RO 50%

The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards

Sukai Huang, Shu-Wei Liu, Nir Lipovetzky, Trevor Cohn

机构 * Google DeepMind(谷歌DeepMind)

专题命中 多模态Agent :multimodal(abstract)

Comments accepted by PRL Workshop Series @ ICAPS 2025. 11 main body pages, 21 appendix pages

详情

展开后加载摘要…

URL PDF HTML 收藏