arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9800 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9132 篇

2605.06247 2026-05-08 cs.RO 79%

CKT-WAM: Parameter-Efficient Context Knowledge Transfer Between World Action Models

CKT-WAM: 在世界动作模型之间高效传递上下文知识

Yuhua Jiang, Yijun Guo, Hongbing Yang, Guojun Lei, Nuo Chen, Yinuo Zhang, Shaoqiang Yan, Bo Lin, Feifei Gao, Biqing Qi

机构 * Tsinghua University(清华大学) LivsynRobotics Shanghai AI Laboratory(上海人工智能实验室)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

AI总结 CKT-WAM通过文本嵌入空间中的紧凑上下文传递教师WAM知识,提升零样本泛化能力,在LIBERO-Plus上达到86.1%的成功率,且在现实任务中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19043 2026-05-08 cs.AI 79%

Learning Lifted Action Models from Unsupervised Visual Traces

从无监督视觉轨迹中学习提升的动作模型

Kai Xi, Stephen Gould, Sylvie Thiébaux

机构 * School of Computing, The Australian National University(澳大利亚国立大学计算机学院)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

AI总结 本文提出一种深度学习框架,通过无监督视觉轨迹学习状态预测、动作预测及提升的动作模型,并引入混合整数线性规划防止预测崩溃,提升模型在多领域中的全局一致性。

Comments Accepted to the 36th International Conference on Automated Planning and Scheduling (ICAPS-26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25859 2026-05-05 cs.RO 79%

Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models

特权预见蒸馏:世界动作模型的零成本未来修正

Pengcheng Fang, Hongli Chen, Xiaohao Cai

机构 * The University of Southampton(南安普顿大学) The University of Queensland(昆士兰大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

AI总结 本文提出特权预见蒸馏,通过将未来信息作为可压缩的修正项进行蒸馏,提升世界动作模型在无未来信息下的表现,实验证明其在LIBERO和RoboTwin任务中有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26489 2026-04-30 cs.LG cs.IR 79%

Understanding DNNs in Feature Interaction Models: A Dimensional Collapse Perspective

理解特征交互模型中的DNN:从维度崩溃视角出发

Jiancheng Wang, Mingjia Yin, Hao Wang, Enhong Chen

机构 * University of Science Technology of China \& State Key Laboratory of Cognitive Intelligence Hefei China Technology of China \& State Key Laboratory of Cognitive Intelligence

专题命中 VLA模型 :action model(title,abstract);分类 cs.LG

AI总结 本文从维度鲁棒性角度探讨DNN在特征交互模型中的有效性,通过实验验证并分析DNN缓解嵌入维度崩溃的机制。

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26259 2026-04-16 cs.IR cs.AI cs.CL 79%

Working Notes on Late Interaction Dynamics: Analyzing Targeted Behaviors of Late Interaction Models

晚期交互动力学工作笔记:分析晚期交互模型的目标行为

Antoine Edy, Max Conti, Quentin Macé

机构 * Illuin Technology(Illuin技术)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

AI总结 本文研究了晚期交互检索中长度偏差和最大相似性操作符外的相似性分布,揭示了因果模型和双向模型的实际性能瓶颈。

Comments Accepted at The 1st Late Interaction Workshop (LIR) @ ECIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08685 2026-04-13 cs.AI 79%

RAMP: Hybrid DRL for Online Learning of Numeric Action Models

RAMP:用于在线学习数字动作模型的混合深度强化学习

Yarin Benyamin, Argaman Mordoch, Shahaf S. Shperberg, Roni Stern

机构 * Ben-Gurion University of the Negev(内盖夫本-古里安大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

AI总结 RAMP通过环境交互在线学习数字规划动作模型,结合深度强化学习策略、动作模型学习和规划,形成正反馈循环提升求解能力和计划质量。

Comments Accepted as a workshop paper at the Adaptive and Learning Agents (ALA) Workshop at AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01570 2026-04-03 cs.RO 79%

Boosting Vision-Language-Action Finetuning with Feasible Action Neighborhood Prior

通过可行动作邻域先验提升视觉-语言-动作微调

Haochen Niu, Kanyu Zhang, Shuyu Yin, Qinghai Guo, Peilin Liu, Fei Wen

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Technologies(华为技术有限公司)

专题命中 VLA模型 :vision-language-action(title);VLA(abstract);分类 cs.RO

AI总结 本文提出FAN引导正则化方法,通过调整模型输出分布以匹配FAN几何特性,提升VLA适应的样本效率和泛化能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26360 2026-03-30 cs.RO 79%

Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and Accurate

Realtime-VLA V2:学习以快速、平滑和准确地运行VLAs

Chen Yang, Yucheng Hu, Yunchao Ma, Yunhuan Yang, Jing Tan, Haoqiang Fan

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO

AI总结 本文提出Realtime-VLA V2,通过校准、规划与控制及学习方法,实现机器人在真实任务中高速、高精度和高灵活性的VLAs运行。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05815 2026-03-09 cs.RO 79%

Hierarchical Latent Action Model

层次化潜在动作模型

Hanjung Kim, Lerrel Pinto, Seon Joo Kim

机构 * Yonsei University(延世大学) New York University(纽约大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

AI总结 HiLAM通过建模长期时间信息,从无动作视频中发现高层次潜在技能,提升了动态技能发现的鲁棒性。

Comments ICLR 2026 Workshop - 2nd Workshop on World Models: Understanding, Modelling and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08946 2026-03-04 q-bio.BM cs.LG 79%

Physically Valid Biomolecular Interaction Modeling with Gauss-Seidel Projection

基于高斯-塞德尔投影的物理有效生物分子相互作用建模

Siyuan Chen, Minghao Guo, Caoliwen Wang, Anka He Chen, Yikun Zhang, Jingjing Chai, Yin Yang, Wojciech Matusik, Peter Yichen Chen

机构 * University of British Columbia(不列颠哥伦比亚大学) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) NVIDIA(英伟达) Peking University(北京大学) Foundry Biosciences(Foundry 生物科技) University of Utah(犹他大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.LG

AI总结 本文提出了一种基于高斯-塞德尔投影的生物分子相互作用建模方法,通过强制物理有效性约束提升结构准确性,实现高效且稳定的去噪过程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01700 2026-03-03 cs.RO 79%

TacMamba: A Tactile History Compression Adapter Bridging Fast Reflexes and Slow VLA Reasoning

TacMamba: 一种连接快速反射与慢速VLA推理的触觉历史压缩适配器

Zhenan Wang, Yanzhe Wang, Meixuan Ren, Peng Li, Yang Liu, Yifei Nie, Limin Long, Yun Ye, Xiaofeng Wang, Zhen Zhu, Huixu Dong

机构 * Zhejiang University(浙江大学) GigaAI Harbin Institute of Technology(哈尔滨工业大学)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO

AI总结 TacMamba通过触觉历史压缩器和双阶段训练策略,实现高频率触觉数据与低频视觉推理的高效融合,提升实时任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00615 2026-03-03 cs.RO 79%

TGM-VLA: Task-Guided Mixup for Sampling-Efficient and Robust Robotic Manipulation

TGM-VLA:基于任务的混合学习用于高效且鲁棒的机器人操作

Fanqi Pu, Lei Jiang, Wenming Yang

机构 * Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(清华大学深圳国际研究生院) The National and Local Co-Build Humanoid Robotics Innovation Center(国家级与地方共建人形机器人创新中心)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO

AI总结 TGM-VLA通过优化关键帧采样策略和引入颜色反转投影模块,提升机器人操作任务的效率和鲁棒性。

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20577 2026-02-25 cs.CV 79%

Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion

高效的端到端自动驾驶:通过掩码视觉-语言-动作扩散

Jiaru Zhang, Manav Gagvani, Can Cui, Juntong Peng, Ruqi Zhang, Ziran Wang

机构 * Institute for Physical Artificial Intelligence (IPAI), Purdue University(物理人工智能研究所(IPAI)、普渡大学) College of Engineering, Purdue University(工程学院、普渡大学) Department of Computer Science, Purdue University(计算机科学系、普渡大学)

专题命中 VLA模型 :vision-language-action(title,abstract);分类 cs.CV

AI总结 MVLAD-AD通过掩码视觉-语言-动作扩散模型,提升自动驾驶的效率与规划精度,实现高效且可解释的端到端自动驾驶。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20383 2026-01-29 cs.CV 79%

HINT: Hierarchical Interaction Modeling for Autoregressive Multi-Human Motion Generation

HINT: 用于自回归多人体运动生成的分层交互建模

Mengge Liu, Yan Di, Gu Wang, Yun Qu, Dekai Zhu, Yanyan Li, Xiangyang Ji

机构 * Tsinghua University(清华大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

AI总结 HINT通过分层交互建模提出了一种自回归框架,实现多人体运动生成,提升了交互建模的精度和长序列的连贯性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19005 2026-01-28 cs.IR cs.LG stat.ML 79%

Recommending Composite Items Using Multi-Level Preference Information: A Joint Interaction Modeling Approach

利用多级偏好信息推荐复合物品:一种联合交互建模方法

Xuan Bi, Yaqiong Wang, Gediminas Adomavicius, Shawn Curley

机构 * Carlson School of Management, University of Minnesota(明尼苏达大学卡尔森管理学院) Leavey School of Business, Santa Clara University(圣克拉拉大学莱维商学院)

专题命中 VLA模型 :action model(title,abstract);分类 cs.LG

AI总结 本文提出JIMA方法,通过联合交互建模利用多级偏好信息,提升复合物品推荐的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11620 2025-12-15 cs.RO cs.SY eess.SY 79%

Architecting Large Action Models for Human-in-the-Loop Intelligent Robots

为具有人类在环的智能机器人构建大型动作模型

Kanisorn Sangchai, Methasit Boonpun, Withawin Kraipetchara, Paulo Garcia

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

AI总结 本文提出通过整合符号方法与现成模型构建可验证的神经符号智能机器人动作模型,以提升可靠性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12405 2025-11-18 cs.CV 79%

VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving

Hyunki Seong, Seongwoo Moon, Hojin Ahn, Jehun Kang, David Hyunchul Shim

机构 * School of Electrical Engineering, KAIST(韩国科学技术院电子工程学院)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.CV

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13397 2025-11-13 cs.CV 79%

Trustworthy Pedestrian Trajectory Prediction via Pattern-Aware Interaction Modeling

Kaiyuan Zhai, Juan Chen, Chao Wang, Zeyi Xu, Guoming Tang

机构 * Shanghai University(上海大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27383 2025-11-03 cs.AI 79%

Realistic pedestrian-driver interaction modelling using multi-agent RL with human perceptual-motor constraints

Yueyang Wang, Mehmet Dogar, Gustav Markkula

机构 * Institute for Transport Studies(交通研究所在) School of Computer Science(计算机科学学院)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14630 2025-09-19 cs.RO 79%

Toward Embodiment Equivariant Vision-Language-Action Policy

Anzhe Chen, Yifei Yang, Zhenjie Zhu, Kechun Xu, Zhongxiang Zhou, Rong Xiong, Yue Wang

专题命中 VLA模型 :vision-language-action(title,abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15885 2025-08-26 cs.AI 79%

VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making

Zuojin Tang, Bin Hu, Chenyang Zhao, De Ma, Gang Pan, Bin Liu

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Home Robotics Lab, E-surfing Digital Life Technology Co., Ltd., China Telecom(E-surfing数字生活技术有限公司) Zhejiang Lab(浙江实验室) Trinity College Dublin(都柏林大学)

专题命中 VLA模型 :action model(title);VLA(abstract);分类 cs.AI

Comments Accepted by EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00920 2025-08-13 physics.chem-ph cs.LG 79%

Uni-Mol3: A Multi-Molecular Foundation Model for Advancing Organic Reaction Modeling

Lirong Wu, Junjie Wang, Zhifeng Gao, Xiaohong Ji, Rong Zhu, Xinyu Li, Linfeng Zhang, Guolin Ke, Weinan E

机构 * AI for Science Institute(人工智能科学研究院) DP Technology(DP技术) College of Chemistry and Molecular Engineering(化学与分子工程学院) School of Mathematical Sciences(数学科学学院) Center for Machine Learning Research(机器学习研究中心)

专题命中 VLA模型 :action model(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18775 2025-07-28 cs.AI cs.SE 79%

Initial Steps in Integrating Large Reasoning and Action Models for Service Composition

Ilche Georgievski, Marco Aiello

机构 * University of Stuttgart(斯图加特大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

Comments 16 pages, 3 figures, 19th Symposium and Summer School on Service-Oriented Computing (SummerSOC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23121 2025-04-01 cs.CV 79%

Efficient Explicit Joint-level Interaction Modeling with Mamba for Text-guided HOI Generation

Guohong Huang, Ling-An Zeng, Zexin Zheng, Shengbo Gu, Wei-Shi Zheng

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

Comments Accepted to ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13006 2025-02-19 cs.AI 79%

Integrating Reinforcement Learning, Action Model Learning, and Numeric Planning for Tackling Complex Tasks

Yarin Benyamin, Argaman Mordoch, Shahaf S. Shperberg, Roni Stern

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10047 2025-01-14 cs.AI 79%

Large Action Models: From Inception to Implementation

Lu Wang, Fangkai Yang, Chaoyun Zhang, Junting Lu, Jiaxu Qian, Shilin He, Pu Zhao, Bo Qiao, Ray Huang, Si Qin, Qisheng Su, Jiayi Ye, Yudi Zhang, Jian-Guang Lou, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

Comments 25pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09631 2024-12-04 cs.AI 79%

Action Model Learning with Guarantees

Diego Aineto, Enrico Scala

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

Journal ref Proceedings of the International Conference on Principles of Knowledge Representation and Reasoning, 21(1), 801-811, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03307 2024-11-25 cs.RO cs.SY eess.SY 79%

Bi-level Trajectory Optimization on Uneven Terrains with Differentiable Wheel-Terrain Interaction Model

Amith Manoharan, Aditya Sharma, Himani Belsare, Kaustab Pal, K. Madhava Krishna, Arun Kumar Singh

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

Comments 8 pages, 7 figures, submitted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08848 2024-11-19 cs.RO 79%

Learning Spatial Bimanual Action Models Based on Affordance Regions and Human Demonstrations

Björn S. Plonka, Christian Dreher, Andre Meixner, Rainer Kartmann, Tamim Asfour

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

Comments 8 pages, accepted for publication at Humanoids 2024 - Copyright IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23218 2024-10-31 cs.CL cs.CV cs.HC 79%

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, Yu Qiao

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏