arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2783 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2783 篇

2602.03175 2026-02-23 cs.LG 50%

Probe-then-Commit Multi-Objective Bandits: Theoretical Benefits of Limited Multi-Arm Feedback

探查-然后承诺多目标老虎机:有限多臂反馈的理论优势

Ming Shi

机构 * Department of Electrical Engineering, University at Buffalo(电气工程系,布法罗大学)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出PtC-P-UCB算法,通过有限探查提升多目标老虎机的性能,实现$1/\sqrt{q}$的加速效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13404 2026-02-17 astro-ph.EP physics.pop-ph 50%

The Interplanetary Habitable Zone

星际宜居带

Caleb Scharf

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出了一种评估星际宜居带的多模式指标,并通过基于代理的模型探讨了星际生命扩散与资源利用之间的平衡。

Comments 38 pages, 14 color figures, submitted to The Astrobiology Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10285 2026-02-17 cs.RO 50%

Adaptive Time Step Flow Matching for Autonomous Driving Motion Planning

自适应时间步长流匹配用于自动驾驶运动规划

Ananya Trivedi, Anjian Li, Mohamed Elnoor, Yusuf Umut Ciftci, Avinash Singh, Jovin D'sa, Sangjae Bae, David Isele, Taskin Padir, Faizan M. Tariq

机构 * HRI Honda Research Institute(本田研究院) Northeastern University(东北大学) Princeton University(普林斯顿大学) University of Maryland(马里兰大学) Stanford University(斯坦福大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种基于条件流匹配的自适应时间步长框架,用于实时自动驾驶轨迹规划,通过在线调整推理步骤数和轨迹后处理提升性能。

Comments Accepted to Intelligent Vehicles Symposium 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12565 2026-02-16 cs.HC 50%

GatheringSense: AI-Generated Imagery and Embodied Experiences for Understanding Literati Gatherings

GatheringSense: 基于人工智能生成影像与具身体验的文人聚会理解

You Zhou, Bingyuan Wang, Hongcheng Guo, Rui Cao, Zeyu Wang

专题命中 多模态Agent :multimodal(abstract)

AI总结 GatheringSense通过AI生成影像与具身体验相结合,探索文人聚会的文化理解与共鸣,提出双路径框架并提供设计启示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12072 2026-02-13 stat.AP 50%

Enhanced Forest Inventories for Habitat Mapping: A Case Study in the Sierra Nevada Mountains of California

增强的森林调查用于栖息地制图:加利福尼亚州内华达山脉的案例研究

Maxime Turgeon, Michael Kieser, Dwight Wolfe, Bruce MacArthur

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出增强森林调查方法,通过整合遥感数据与地面调查,用于高分辨率野生动物栖息地制图,为森林管理和生态保护提供空间明确的工具。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10109 2026-02-11 cs.RO 50%

ST4VLA: Spatially Guided Training for Vision-Language-Action Models

ST4VLA:基于空间引导的视觉-语言-动作模型训练

Jinhui Ye, Fangjing Wang, Ning Gao, Junqiu Yu, Yangkun Zhu, Bin Wang, Jinyu Zhang, Weiyang Jin, Yanwei Fu, Feng Zheng, Yilun Chen, Jiangmiao Pang

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Hong Kong University of Science and Technology(香港科技大学) Southern University of Science and Technology(南方科技大学) Fudan University(复旦大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 ST4VLA通过空间引导训练提升视觉-语言-动作模型的性能,实现对机器人任务的更稳健和可泛化的学习。

Comments Spatially Training for VLA, Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10054 2026-02-11 cs.HC 50%

AIDED: Augmenting Interior Design with Human Experience Data for Designer-AI Co-Design

AIDED: 通过人类经验数据增强室内设计以实现设计师-人工智能协同设计

Yang Chen Lin, Chen-Ying Chen, Kai-Hsin Hou, Hung-Yu Chen, Po-Chih Kuo

专题命中 多模态Agent :multimodal(abstract)

AI总结 AIDED通过整合人类经验数据提升室内设计的创造力和客户互动,探索不同数据模态对设计结果的影响,并提出未来AI工具的设计方向。

Comments 29 pages, 14 figures, Accepted to CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01100 2026-02-10 cs.RO 50%

StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating

StreamVLA: 通过完成状态门控打破原因-行动循环

Tongqing Chen, Hang Wu, Jiasen Wang, Xiaotao Li, Lu Fang

机构 * Tsinghua University(清华大学) Technical University of Munich(慕尼黑技术大学) Shanghai University(上海大学) Wuhan University(武汉大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 StreamVLA通过完成状态门控机制,实现高效机器人操作,减少推理延迟并提升任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23189 2026-02-10 eess.SY cs.SY 50%

The Dawn of Agentic EDA: A Survey of Autonomous Digital Chip Design

代理EDA的黎明:自主数字芯片设计的综述

Zelin Zang, Yuhang Song, Aili Wang, Bingo Wing-Kuen Ling, Qi Sun, Zhen Lei, Fuji Yang, Cheng Zhuo, Jiebo Luo

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出代理EDA的系统框架,通过认知堆栈分析前端和后端设计流程,强调神经符号优化和可信度提升的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11510 2026-02-10 stat.OT stat.CO 50%

Here Be Dragons: Bimodal posteriors arise from numerical integration error in longitudinal models

此处有龙:纵向模型中双峰后验分布源于数值积分误差

Tess O'Brien, Matthew T. Moores, David Warton, Daniel Falster

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文研究了纵向模型中数值积分误差导致的双峰后验分布问题,提出通过合适方法避免双峰性,对贝叶斯反演方法有重要影响。

Comments 33 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06966 2026-02-10 cs.RO 50%

Embodied Intelligence for Flexible Manufacturing: A Survey

具身智能在柔性制造中的应用:综述

Kai Xu, Hang Zhao, Ruizhen Hu, Min Yang, Hao Liu, Hui Zhang, Haibin Yu

机构 * National University of Defense Technology(中国人民解放军国防科技大学) Wuhan University(武汉大学) Shenzhen University(深圳大学) Hunan University(湖南大学) Shenyang Institute of Automation Chinese Academy of Sciences(中国科学院沈阳自动化研究所)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文综述了具身智能在柔性制造中的应用,分析了感知、控制和决策三个层面的技术挑战,并提出三阶段进化模型以指导其发展。

Comments in chinese language. ROBOT

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06030 2026-02-10 cs.MA cs.LG 50%

PhysicsAgentABM: Physics-Guided Generative Agent-Based Modeling

PhysicsAgentABM: 基于物理的生成基于主体的建模

Kavana Venkatesh, Yinhan He, Jundong Li, Jiaming Cui

机构 * University of Virginia(弗吉尼亚大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 PhysicsAgentABM通过结合符号推理与神经网络,实现基于物理的生成式主体建模,提升大规模模拟的准确性和校准性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00260 2026-02-06 eess.SP 50%

Sensor Insoles: A Review

智能传感鞋垫:综述

Bastian Latsch, Felix Herbst, Mark Suppelt, Julian Seiler, Stephan Schaumann, Sven Suppelt, Alexander A. Altmann, Martin Grimmer, and Mario Kupnik

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文综述了智能传感鞋垫在足部压力测量中的应用,分析了现有技术的局限性,并提出未来多模态传感器和多轴传感的发展方向。

Comments 20 pages, 8 figures, review article published in IEEE Sensors Journal

Journal ref IEEE Sensors Journal, vol. 26, no. 3, pp. 3577-3596, Dec. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03153 2026-02-04 cs.RO 50%

When Attention Betrays: Erasing Backdoor Attacks in Robotic Policies by Reconstructing Visual Tokens

当注意力背叛时:通过重建视觉标记在机器人策略中消除后门攻击

Xuetao Li, Pinhan Fu, Wenke Huang, Nengyuan Pan, Songhua Yang, Kaiyan Zhao, Guancheng Wan, Mengde Li, Jifeng Xuan, Miao Li

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Faculty of Artificial Intelligence, Hubei University(湖北大学人工智能学院) Institute of Technological Sciences, Wuhan University(武汉大学技术科学研究院) School of Robotics, Wuhan University(武汉大学机器人学院)

专题命中 多模态Agent :multimodal(abstract)

AI总结 Bera通过重建视觉标记消除机器人策略中的后门攻击,无需重新训练模型即可有效防御后门风险。

Comments ICRA2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02959 2026-02-04 cs.LG cs.SY eess.SY 50%

Human-Centric Traffic Signal Control for Equity: A Multi-Agent Action Branching Deep Reinforcement Learning Approach

以人为中心的交通信号控制以实现公平性:一种多智能体动作分支深度强化学习方法

Xiaocai Zhang, Neema Nassir, Lok Sang Chan, Milad Haghani

机构 * Department of Infrastructure Engineering, Faculty of Engineering and Information Technology, The University of Melbourne, VIC 3010, Australia(工程与信息技术学院基础设施工程系,墨尔本大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种以人为中心的多智能体动作分支深度强化学习方法,通过分解交通走廊控制并优化旅行者公平性,有效减少受延迟旅行者数量,提升交通信号系统的公平性和可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12311 2026-02-04 cs.RO 50%

Scene-Adaptive Motion Planning with Explicit Mixture of Experts and Interaction-Oriented Optimization

场景自适应运动规划与显式专家混合及交互导向优化

Hongbiao Zhu, Liulong Ma, Xian Wu, Xin Deng, Xiaoyao Liang

机构 * Automotive New Technology Research Institute, BYD Company Limited(比亚迪公司自动驾驶新技术研究院)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出EMoE-Planner,通过显式专家混合和交互导向优化,解决复杂城市环境中自主驾驶轨迹规划的多模态、场景多样性和环境交互问题,实现超越传统规则算法的性能。

Comments Main text 10 pages with 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08952 2026-02-03 cs.LG 50%

When LLM Agents Meet Graph Optimization: An Automated Data Quality Improvement Approach

当大语言模型代理遇见图优化:一种自动化数据质量改进方法

Zhihan Zhang, Xunkai Li, Yilong Zuo, Yanzhe Wen, Zhaoxin Fan, Zhenjun Li, Bing Zhou, Rong-Hua Li, Guoren Wang

机构 * Shenzhen Institute of Technology(深圳理工大学)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 LAGA是一种统一的多代理框架,通过协调多模态优化全面提升文本属性图的数据质量,验证了数据驱动的质量优化在可靠分析中的重要性。

Comments 12 pages, 10figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22431 2026-02-03 cs.SE 50%

TreeMind: Automatically Reproducing Android Bug Reports via LLM-empowered Monte Carlo Tree Search

TreeMind: 通过LLM赋能的蒙特卡洛树搜索自动重现Android bug报告

Zhengyu Chen, Zhaoyi Meng, Wenxiang Zhao, Wansen Wang, Wenchao Huang, Jie Cui, Hong Zhong, Yan Xiong

专题命中 多模态Agent :multi-modal(abstract)

AI总结 TreeMind结合LLM与MCTS算法,通过语义推理和反馈导航,高效重现Android bug报告。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18745 2026-02-02 cond-mat.mtrl-sci 50%

Spin defects in hexagonal boron nitride as two-dimensional strain sensors

六方氮化硼中的自旋缺陷作为二维应变传感器

Z. Mu, Z. Zhang, J. Fraunié, C. Robert, G. Seine, B. Gil, G. Cassabois, V. Jacques

专题命中 多模态Agent :multimodal(abstract)

AI总结 六方氮化硼中的硼空位色心被证实可作为高精度二维应变传感器,用于测量多层hBN晶片在应力下的拉曼模式位移,为范德瓦尔异质结构的应变计量提供新工具。

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25889 2026-01-30 cs.LG 50%

$π_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

$π_\ exttt{RL}$: 流基于视觉-语言-动作模型的在线强化学习微调

Kang Chen, Zhihao Liu, Tonghe Zhang, Zhen Guo, Si Xu, Hao Lin, Hongzhi Zang, Xiang Li, Quanlu Zhang, Zhaofei Yu, Guoliang Fan, Tiejun Huang, Yu Wang, Chao Yu

机构 * Tsinghua University(清华大学) Peking University(北京大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Carnegie Mellon University(卡内基梅隆大学) Infinigence AI Zhongguancun Academy(中关村学院)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出$π_\ exttt{RL}$方法,通过流噪声和流SDE技术,解决大规模流基于VLA模型中强化学习微调的挑战,提升模型在分布内和分布外任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18492 2026-01-27 cs.RO 50%

DV-VLN: Dual Verification for Reliable LLM-Based Vision-and-Language Navigation

DV-VLN:基于可靠大语言模型的视觉-语言导航的双重验证

Zijun Li, Shijie Li, Zhenxi Zhang, Bin Li, Shoujun Zhou

机构 * Robotics Engineering Program, College of Engineering, Zhejiang Normal University(浙江师范大学工程学院机器人工程专业) Shenzhen Institutes of Advanced Technology (SIAT), Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Department of Health Technology and Informatics, The Hong Kong Polytechnic University(香港理工大学健康科技与信息技术系)

专题命中 多模态Agent :cross-modal(abstract)

AI总结 DV-VLN通过生成-验证范式提升大语言模型在视觉-语言导航中的可靠性,通过双重验证机制提高导航决策的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17548 2026-01-27 cs.CR 50%

Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems

针对代理编码助手的提示注入攻击:对技能、工具和协议生态系统漏洞的系统分析

Narek Maloyan, Dmitry Namiot

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文系统分析了代理编码助手的提示注入攻击漏洞,提出三维分类法并揭示防御不足,强调需架构级防护而非碎片化措施。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13801 2026-01-21 cs.RO 50%

HoverAI: An Embodied Aerial Agent for Natural Human-Drone Interaction

HoverAI: 一种用于自然人-无人机交互的具身空中代理

Yuhua Jin, Nikita Kuzmin, Georgii Demianchuk, Mariya Lezina, Fawad Mehboob, Issatay Tokmurziyev, Miguel Altamirano Cabrera, Muhammad Ahsan Mustafa, Dzmitry Tsetserukou

机构 * Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Skolkovo Institute of Science and Technology(斯克尔科沃信息科技研究所)

专题命中 多模态Agent :multimodal(abstract)

AI总结 HoverAI通过结合无人机移动、视觉投影和对话式AI,实现了人-无人机自然交互的具身代理,提升了空间感知与社交响应能力。

Comments This paper has been accepted for publication at LBR HRI 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12993 2026-01-21 cs.RO 50%

Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization

Being-H0.5:面向跨躯体泛化的以人为本的机器人学习规模化

Hao Luo, Ye Wang, Wanpeng Zhang, Sipeng Zheng, Ziheng Xi, Chaoyi Xu, Haiweng Xu, Haoqi Yuan, Chi Zhang, Yiqing Wang, Yicheng Feng, Zongqing Lu

机构 * BeingBeyond Team(BeingBeyond 团队)

专题命中 多模态Agent :multimodal(abstract)

AI总结 Being-H0.5通过以人为本的学习范式和统一动作空间,实现了在不同机器人平台间的跨躯体泛化,结合混合流框架和流形保持门控,达到模拟和现实中的高性能表现。

Comments 44 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01147 2026-01-21 cs.RO 50%

Astra: Efficient Transformer Architecture and Contrastive Dynamics Learning for Embodied Instruction Following

Astra:面向具身指令跟随的高效Transformer架构与对比动态学习

Yueen Ma, Dafeng Chi, Shiguang Wu, Yuecheng Liu, Yuzheng Zhuang, Irwin King

机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(计算机科学与工程系,香港中文大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态Agent :multimodal(abstract)

AI总结 Astra通过引入轨迹注意力和对比动态学习目标,提升了具身指令跟随任务中多模态序列处理的效率与准确性。

Comments Accepted to EMNLP 2025 (main). Published version: https://aclanthology.org/2025.emnlp-main.688/ Code available at: https://github.com/yueen-ma/Astra

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10208 2026-01-16 cs.RO 50%

Terrain-Adaptive Mobile 3D Printing with Hierarchical Control

地形自适应移动三维打印与分层控制

Shuangshan Nors Li, J. Nathan Kutz

机构 * Department of Electrical and Computer Engineering, University of Washington, USA(电气与计算机工程系,华盛顿大学) Department of Applied Mathematics, University of Washington, USA(应用数学系,华盛顿大学)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本研究提出了一种结合人工智能与分层控制的地形自适应移动三维打印框架,通过多传感器融合和闭环控制实现高精度打印与移动性平衡。

Comments Submitted to the 43rd International Symposium on Automation and Robotics in Construction (ISARC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09937 2026-01-16 cs.HC cs.IR 50%

From SERPs to Agents: A Platform for Comparative Studies of Information Interaction

从搜索结果页面到代理:一个用于信息交互比较研究的平台

Saber Zerhoudi, Michael Granitzer

专题命中 多模态Agent :multi-modal(abstract)

AI总结 UXLab是一个开源平台,用于比较信息交互系统,通过无代码实验设计支持多模态交互研究。

Journal ref Proceedings of the 2026 ACM SIGIR Conference on Human Information Interaction and Retrieval (CHIIR '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04267 2026-01-16 physics.soc-ph cs.MA 50%

Information Theoretic Optimal Surveillance for Epidemic Prevalence in Networks

信息论最优监控以估计网络中的疫情流行率

Ritwick Mishra, Abhijin Adiga, Madhav Marathe, S. S. Ravi, Ravi Tandon, Anil Vullikanti

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出TESTPREV问题,通过最大化互信息选择网络节点以优化疫情流行率估计,提出GREEDYMI策略在IC模型下提升互信息和减少方差。

Comments 25 pages; Added acknowledgments; In Proceedings of the AAAI 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07252 2026-01-13 cs.MA 50%

SwarmFoam: An OpenFOAM Multi-Agent System Based on Multiple Types of Large Language Models

SwarmFoam: 基于多种大语言模型的开放FOAM多智能体系统

Chunwei Yang, Yankai Wang, Jianxiang Tang, Haojie Qu, Ziqiang Zou, YuLiu, Chunrui Deng, Zhifang Qiu, Ming Ding

专题命中 多模态Agent :multi-modal(abstract)

AI总结 SwarmFoam通过整合多模态感知与智能校正技术,提升复杂CFD模拟的适应性与准确性。

Comments 26 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06920 2026-01-13 cs.NE cs.MA 50%

Calibrating Agent-Based Financial Markets Simulators with Pretrainable Automatic Posterior Transformation-Based Surrogates

基于可预训练的自动后验变换的代理基于金融市场模拟器校准

Boquan Jiang, Zhenhua Yang, Chenkai Wang, Muyao Zhong, Heping Fang, Peng Yang

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出ANTR方法,通过可预训练的神经密度估计器和自适应信任区域策略,提升代理基于金融市场模拟器的校准精度和效率。

Comments 32 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏