arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2783 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2783 篇

2603.20355 2026-03-24 eess.IV 50%

CaroTo: A Tool for Fast Comprehensive Analysis of Carotid Artery Stenosis in 4D PC- and 3D BB-MRI Data

CaroTo:一种用于快速全面分析4D PC-和3D BB-MRI数据颈动脉狭窄的工具

Hinrich Rahlfs, Markus Hüllebrand, Sebastian Schmitter, Jonathan Andrae, Christoph Strecker, Andreas Harloff, Anja Hennemuth

专题命中 多模态Agent :multimodal(abstract)

AI总结 CaroTo工具通过多模态和多维分割、生物标志物提取和可视化,实现颈动脉动脉粥样斑块的标准化评估,提升颈动脉狭窄分析的精度和一致性。

Comments VCBM 2024, Poster Honorable Mention

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19858 2026-03-23 cs.RO cs.MA 50%

Beyond detection: cooperative multi-agent reasoning for rapid onboard EO crisis response

超越检测:用于快速在轨遥感危机响应的协作多智能体推理

Alejandro D. Mousist, Pedro Delgado de Robles Martín, Raquel Lladró Climent, Julian Cobos Aparicio

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种分层多智能体架构,用于在轨遥感处理,在资源和带宽受限条件下,通过协调专用AI智能体实现互补多模态观测的利用,减少计算开销并保持决策一致性。

Comments Accepted for presentation at the ESA's 4S Symposium 2026 Conference (see https://atpi.eventsair.com/4s-symposium-2026/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19555 2026-03-23 astro-ph.IM 50%

SpecZoo: An AI-Powered Platform for Spectral Analysis and Visualization in Science and Education

SpecZoo:一个基于AI的天文光谱分析与可视化平台

Yuan-Hao Pu, Guo-Hong Lei, Yang Xu, Xun-Zhou Chen, Hai-Jun Tian

专题命中 多模态Agent :multi-modal(abstract)

AI总结 SpecZoo平台利用人工智能技术,整合现代信息技术和机器学习,提升光谱数据处理效率,支持光谱可视化、自动分类、参数测量及多波段数据融合,应用于LAMOST、SDSS等重大项目,并促进天文与数据科学的教育融合。

Comments 19 pages, 11 figures, 2 tables, published in the journal of 'universe' (see the special issue: https://www.mdpi.com/journal/universe/special_issues/77CGKMGC3Q)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17634 2026-03-19 eess.SY cs.SY 50%

Hierarchical Decision-Making under Uncertainty: A Hybrid MDP and Chance-Constrained MPC Approach

在不确定性下的分层决策:一种混合MDP和机会约束MPC方法

Siyuan Li, Chengyuan Liu, Wen-Hua Chen

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出一种混合MDP和机会约束MPC的分层决策框架,用于自动驾驶中处理不确定性,通过多模态预测与安全约束实现 maneuver 选择与动态可行性的统一,验证了其在高速公路和城市环境中的安全性和效率。

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21676 2026-03-18 cs.RO cs.NI 50%

Real-World Deployment of Cloud-based Autonomous Mobility Systems for Outdoor and Indoor Environments

云原生自主移动系统在户外和室内环境中的实际部署

Yufeng Yang, Minghao Ning, Keqi Shu, Aladdin Saleh, Ehsan Hashemi, Amir Khajepour

机构 * Department of Mechanical and Mechatronics Engineering, University of Waterloo(滑铁卢大学机械与机电工程系) Technology Partnerships and Innovations, Rogers Communications, Canada Inc.(罗杰斯通讯加拿大有限公司技术伙伴关系与创新部) Mechanical Engineering Department, University of Alberta(阿尔伯塔大学机械工程系)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出云原生自主移动框架,通过基础设施智能传感与云计算协调提升自主操作能力,实验证明在城市环岛和医院类室内环境中的感知鲁棒性和安全性提升。

Comments This paper has been submitted to IEEE Robotics and Automation Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18373 2026-03-18 cs.RO cs.HC 50%

UGotMe: An Embodied System for Affective Human-Robot Interaction

UGotMe: 一种用于情感人机交互的具身系统

Peizhen Li, Longbing Cao, Xiao-Ming Wu, Xiaohan Yu, Runze Yang

机构 * School of Computing, Macquarie University(麦考瑞大学计算学院) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Department of Automation, Shanghai Jiao Tong University(上海交通大学自动化学院)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出UGotMe系统,解决多对话场景中视觉噪声和实时响应问题,通过去噪策略和高效数据传输提升情感识别能力。

Comments Accepted to the 2025 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08478 2026-03-17 cs.RO cs.LG 50%

STRIDE: Structured Lagrangian and Stochastic Residual Dynamics via Flow Matching

STRIDE: 通过流匹配实现结构化拉格朗日和随机残差动力学

Prakrut Kotecha, Ganga Nair B, Shishir Kolathaya

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出STRIDE框架,结合拉格朗日神经网络与条件流匹配,分离保守刚体动力学与非保守随机交互效应,提升机器人在不确定环境中的预测精度与控制可靠性。

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13840 2026-03-17 cs.MA 50%

ClimateAgents: A Multi-Agent Research Assistant for Social-Climate Dynamics Analysis

ClimateAgents: 一种用于社会-气候动态分析的多智能体研究助手

Shan Shan

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出ClimateAgents,一种多智能体研究助手,通过整合多模态数据检索、统计建模和自动推理,帮助研究者探索社会-环境动态,提升气候分析的适应性和解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12516 2026-03-16 cs.LG physics.flu-dyn 50%

Learning Pore-scale Multiphase Flow from 4D Velocimetry

从4D速度测距学习孔隙尺度多相流

Chunyang Wang, Linqi Zhu, Yuxuan Gu, Robert van der Merwe, Xin Ju, Catherine Spurin, Samuel Krevor, Rex Ying, Tobias Pfaff, Martin J. Blunt, Tom Bultreys, Gege Wen

机构 * Department of Earth Science and Engineering, Imperial College London(帝国理工学院伦敦地球科学与工程系) Department of Geology, Ghent University(根特大学地质系) Department of Energy Science and Engineering, Stanford University(斯坦福大学能源科学与工程系) Department of Computer Science, Yale University(耶鲁大学计算机科学系) NVIDIA

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出一种多模态学习框架,通过4D微速度测距数据直接推断多相孔隙流,结合图网络模拟和3D U-Net,实现快速预测,为地下碳和氢存储提供高效工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11569 2026-03-13 cs.HC 50%

Modeling Sequential Design Actions as Designer Externalization on an Infinite Canvas

将顺序设计动作建模为设计师在无限画布上的外部化

Yejin Yun, Seung Won Lee, Jiin Choi, Kyung Hoon Hyun

专题命中 多模态Agent :multimodal(abstract)

AI总结 研究探讨了AI在无限画布设计中的作用,发现AI通过改变认知分配和工作流程,促进了设计师与AI的协同进化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11476 2026-03-13 cs.LG q-bio.QM 50%

Leveraging Phytolith Research using Artificial Intelligence

利用人工智能进行phytolith研究

Andrés G. Mejía Ramón, Kate Dudgeon, Nina Witteveen, Dolores Piperno, Michael Kloster, Luigi Palopoli, Mónica Moraes R., José M. Capriles, Umberto Lombardo

专题命中 多模态Agent :multimodal(abstract)

AI总结 Sorometry通过人工智能整合2D和3D数据,实现phytolith的高效分析与预测,提升考古和古生态研究的精度与标准化。

Comments 45 pages, 23 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10268 2026-03-12 cs.SE 50%

SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments

SpecOps:一种在真实GUI环境中完全自动化的AI代理测试框架

Syed Yusuf Ahmed, Shiwei Feng, Chanwoo Bae, Calix Barrus Xiangyu Zhang

专题命中 多模态Agent :multimodal(abstract)

AI总结 SpecOps是一种全新的自动化AI代理测试框架,通过专门设计的LLM代理处理四个阶段,实现对真实GUI环境中复杂代理的高效测试与验证。

Comments Accepted to ICSE 2026

Journal ref Proceedings of the IEEE/ACM International Conference on Software Engineering (ICSE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08817 2026-03-11 cs.RO 50%

HMR-1: Hierarchical Massage Robot with Vision-Language-Model for Embodied Healthcare

HMR-1: 基于视觉-语言模型的分层按摩机器人用于具身医疗

Rongtao Xu, Mingming Yu, Xiaofeng Han, Yu Zhang, Kaiyi Hu, Zhe Feng, Zenghuang Fu, Changwei Wang, Weiliang Meng, Xiaopeng Zhang

机构 * The State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, China(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) Spatiotemporal AI, China(时空人工智能,中国) Hangzhou International Innovation Institute, Beihang University, China(杭州国际创新研究院,北航,中国) Georgia Institute of Technology, China(佐治亚理工学院,中国) Key Laboratory of Computing Power Network and Information Security, Ministry of Education(计算功率网络与信息安全重点实验室,教育部;山东省计算机科学中心,齐鲁工业大学(山东省科学院),中国) Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences), China

专题命中 多模态Agent :multimodal(abstract)

AI总结 HMR-1提出基于视觉-语言模型的分层按摩机器人框架,通过构建多模态数据集和微调Qwen-VL模型,实现了具身医疗任务的评估与应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06959 2026-03-11 physics.soc-ph nlin.CD 50%

Discerning media bias within a network of political allies: an analytic condition for disruption by partisans

在政治盟友网络中辨别媒体偏见:一种用于党派破坏的分析条件

Jarra Horstman, Andrew Melatos, Farhad Farokhi

专题命中 多模态Agent :multimodal(abstract)

AI总结 研究在政治盟友网络中辨别媒体偏见的条件,通过概率框架分析党派分子对代理体观点的影响,推导出湍流非收敛与渐近学习的区分条件。

Comments 37 pages, 8 figures

Journal ref Physica A: Statistical Mechanics and its Applications Volume 673, 1 September 2025, 130679

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09587 2026-03-10 cs.RO 50%

GeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation

GeoNav: 通过双尺度地理推理赋能大语言模型实现语言目标的空中导航

Haotian Xu, Yue Hu, Chen Gao, Zhengqiu Zhu, Yong Zhao, Yong Li, Quanjun Yin

机构 * College of Systems Engineering, National University of Defense Technology(系统工程学院,国防科技大学) State Key Laboratory of Digital Intelligent Modeling and Simulation(数字智能建模与仿真国家重点实验室) BNRist, Tsinghua University(清华大学BNRist)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 GeoNav通过双尺度地理推理赋能大语言模型,实现语言目标的空中导航,提升城市场景下的导航成功率和精度。

Comments Published in Pattern Recognition (2026)

Journal ref Pattern Recognition, Volume 177, 113365, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05487 2026-03-06 cs.RO 50%

Observing and Controlling Features in Vision-Language-Action Models

观察和控制视觉-语言-动作模型中的特征

Hugo Buurmeijer, Carmen Amo Alonso, Aiden Swann, Marco Pavone

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出通过特征可观察性和可控性研究,实现对视觉-语言-动作模型的在线适应与行为引导,提升其与用户需求的实时对齐能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05019 2026-03-06 cs.HC 50%

Haptics in Cognition: Disruptor or Enabler of Memory?

触觉在认知中的作用:记忆的破坏者还是促进者?

Bibeg Limbu, Irene-Angelica Chounta

专题命中 多模态Agent :multimodal(abstract)

AI总结 本研究探讨触觉敏感度和运动强度对记忆的影响,发现增加书写压力略微降低即时回忆,但手套使用无明显影响,揭示了身体互动与认知表现之间的复杂关系。

Comments 22 Pages (including references), Book chapter

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04754 2026-03-06 cs.HC 50%

VizCrit: Exploring Strategies for Displaying Computational Feedback in a Visual Design Tool

VizCrit: 探索在视觉设计工具中展示计算反馈的策略

Mingyi Li, Mengyi Chen, Sarah Luo, Yining Cao, Haijun Xia, Maitraye Das, Steven P. Dow, Jane L. E

专题命中 多模态Agent :multi-modal(abstract)

AI总结 VizCrit通过算法问题检测和视觉注释生成,探索在视觉设计工具中实现可操作反馈的策略,发现以解决方案为中心的反馈能提升初学者的创造力感知和设计质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04705 2026-03-06 cs.RO cs.HC 50%

LEGS-POMDP: Language and Gesture-Guided Object Search in Partially Observable Environments

LEGS-POMDP:语言和手势引导的在部分可观测环境中的对象搜索

Ivy Xiao He, Stefanie Tellex, Jason Xinyu Liu

机构 * Brown University(布朗大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 LEGS-POMDP通过整合语言、手势和视觉信息,在部分可观测环境中实现高效的开放世界对象搜索,显著提升了多模态感知和不确定性处理能力。

Comments 10 pages, 8 figures, accepted at ACM/IEEE International Conference on Human-Robot Interaction (HRI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21161 2026-03-06 cs.RO 50%

MarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environments

MarketGen: 一个可扩展的仿真平台,具有自动生成的具身超市环境

Xu Hu, Yiyang Feng, Junran Peng, Jiawei He, Liyi Chen, Wei Sui, Chuanchen Luo, Xucheng Yin, Qing Li, Zhaoxiang Zhang

机构 * The Hong Kong Polytechnic University(香港理工大学) University of Science and Technology Beijing(北京科技大学) NLPR, MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) XYZ Embodied AI(XYZ具身AI) Shandong University(山东大学) Linketic D-Robotics

专题命中 多模态Agent :multi-modal(abstract)

AI总结 MarketGen通过自动生成复杂超市环境,为评估超市代理提供了一个新的基准,加速了复杂商业应用中具身AI的研究。

Comments Project Page: https://xuhu0529.github.io/MarketGen

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04329 2026-03-05 cs.RO 50%

Gaussian Mixture-Based Inverse Perception Contract for Uncertainty-Aware Robot Navigation

基于高斯混合的逆感知合同用于不确定性感知的机器人导航

Bingyao Du, Joonkyung Kim, Yiwei Lyu

机构 * Department of Computer Science, Columbia University(计算机科学系,哥伦比亚大学) Department of Computer Science and Engineering, Texas A&M University(计算机科学与工程系,德克萨斯大学)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出基于高斯混合的逆感知合同,通过联合椭球置信集表示不确定性,提升机器人导航的安全性和适应性。

Comments 8 pages, 5 figures. Accepted to ACC 2026 (American Control Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03121 2026-03-04 cs.SE 50%

RippleGUItester: Change-Aware Exploratory Testing

RippleGUItester: 基于变更的探索性测试

Yanqi Su, Michael Pradel, Chunyang Chen

专题命中 多模态Agent :multimodal(abstract)

AI总结 RippleGUItester通过基于大语言模型的变更影响分析,结合多模态bug检测,发现由代码变更引入的bug,提升测试效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15602 2026-03-04 cs.NE 50%

Estimate Hitting Time by Hitting Probability for Elitist Evolutionary Algorithms

通过击中概率估计精英进化算法的击中时间

Jun He, Siang Yew Chong, Xin Yao

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出通过击中概率的漂移分析方法,用于估计精英进化算法的击中时间,通过引入路径处理多模态适应度景观,简化了击中概率的计算并比较了两种算法的性能。

Journal ref IEEE Transactions on Evolutionary Computation 12 November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00079 2026-03-04 cs.HC 50%

Enhancing the Interpretability of SHAP Values Using Large Language Models

利用大语言模型增强SHAP值的可解释性

Xianlong Zeng, Kewen Zhu

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文利用大语言模型提升SHAP值的可解释性,使非技术用户更易理解模型预测。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01687 2026-03-03 cs.NI 50%

Predictive Importance Sampling Based Coverage Verification for Multi-UAV Trajectory Planning

基于预测重要性采样的多无人机轨迹规划覆盖验证

Snehashish Ghosh, Sasthi C. Ghosh

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出基于预测重要性采样的多无人机轨迹规划覆盖验证方法,通过LSTM-MDN和防御性混合采样提升鲁棒性,实现更高效的覆盖管理。

Comments This article has been submitted to a conference for peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00522 2026-03-03 cs.HC 50%

SIAgent: Spatial Interaction Agent via LLM-powered Eye-Hand Motion Intent Understanding in VR

SIAgent:通过LLM驱动的眼手运动意图理解实现VR中的空间交互代理

Zhimin Wang, Chenyu Gu, Feng Lu

专题命中 多模态Agent :multimodal(abstract)

AI总结 SIAgent通过LLM驱动的意图识别与代理执行,实现VR中基于自然眼手运动的高容忍度交互框架。

Comments Virtual reality, spatial interaction, intent recognition, agent-based execution, large language models

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18676 2026-02-26 cs.HC 50%

MagHeart: Exploring Playful Avatar Co-Creation and Shared Heartbeats for Icebreaking in Hybrid Meetings

MagHeart:探索混合会议中的 playful 虚拟角色共创与共享心跳以促进破冰

Black Sun, Haiyang Xu, Ge Kacy Fu, Liyue Da, Eve Hoggan

专题命中 多模态Agent :multimodal(abstract)

AI总结 MagHeart通过playful虚拟角色共创和共享心跳技术,促进混合会议中的对称破冰,探索远程参与者在会议开始时的物质和感知存在,同时揭示隐私和情境适当性等挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20923 2026-02-25 cs.RO 50%

ParkDiffusion++: Ego Intention Conditioned Joint Multi-Agent Trajectory Prediction for Automated Parking using Diffusion Models

ParkDiffusion++: 基于扩散模型的多智能体轨迹预测用于自动化停车的自我意图条件联合预测

Jiarong Wei, Anna Rehr, Christian Feist, Abhinav Valada

机构 * Department of Computer Science, University of Freiburg(弗赖堡大学计算机科学系) CARIAD SE

专题命中 多模态Agent :multi-modal(abstract)

AI总结 ParkDiffusion++通过联合学习多智能体轨迹预测与自我意图预测,实现自动化停车中的高效假设决策与安全预测。

Comments ICRA 2026 Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19764 2026-02-24 cs.RO 50%

Towards Dexterous Embodied Manipulation via Deep Multi-Sensory Fusion and Sparse Expert Scaling

通过深度多感官融合与稀疏专家扩展实现灵巧的具身体验

Yirui Sun, Guangyu Zhuge, Keliang Liu, Jie Gu, Zhihao xia, Qionglin Ren, Chunxu tian, Zhongxue Ga

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University(智能机器人与先进制造学院,复旦大学) School of Information Science and Engineering, Harbin Institute of Technology(信息科学与工程学院,哈尔滨工业大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 DeMUSE通过深度多感官融合与稀疏专家扩展,实现复杂物理互动的高成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18572 2026-02-24 cs.LG q-fin.ST 50%

Sub-City Real Estate Price Index Forecasting at Weekly Horizons Using Satellite Radar and News Sentiment

基于卫星雷达和新闻情绪的周频次下辖城市房地产价格指数预测

Baris Arat, Hasan Fehmi Ates, Emre Sefer

机构 * Ozyegin University(奥兹根大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文通过结合卫星雷达与新闻情绪数据,实现了下辖城市房地产价格指数的周频预测,显著提升了长期预测的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏