arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9119 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9119 篇

2603.05815 2026-03-09 cs.RO 79%

Hierarchical Latent Action Model

层次化潜在动作模型

Hanjung Kim, Lerrel Pinto, Seon Joo Kim

机构 * Yonsei University(延世大学) New York University(纽约大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

AI总结 HiLAM通过建模长期时间信息,从无动作视频中发现高层次潜在技能,提升了动态技能发现的鲁棒性。

Comments ICLR 2026 Workshop - 2nd Workshop on World Models: Understanding, Modelling and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08946 2026-03-04 q-bio.BM cs.LG 79%

Physically Valid Biomolecular Interaction Modeling with Gauss-Seidel Projection

基于高斯-塞德尔投影的物理有效生物分子相互作用建模

Siyuan Chen, Minghao Guo, Caoliwen Wang, Anka He Chen, Yikun Zhang, Jingjing Chai, Yin Yang, Wojciech Matusik, Peter Yichen Chen

机构 * University of British Columbia(不列颠哥伦比亚大学) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) NVIDIA(英伟达) Peking University(北京大学) Foundry Biosciences(Foundry 生物科技) University of Utah(犹他大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.LG

AI总结 本文提出了一种基于高斯-塞德尔投影的生物分子相互作用建模方法,通过强制物理有效性约束提升结构准确性,实现高效且稳定的去噪过程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01700 2026-03-03 cs.RO 79%

TacMamba: A Tactile History Compression Adapter Bridging Fast Reflexes and Slow VLA Reasoning

TacMamba: 一种连接快速反射与慢速VLA推理的触觉历史压缩适配器

Zhenan Wang, Yanzhe Wang, Meixuan Ren, Peng Li, Yang Liu, Yifei Nie, Limin Long, Yun Ye, Xiaofeng Wang, Zhen Zhu, Huixu Dong

机构 * Zhejiang University(浙江大学) GigaAI Harbin Institute of Technology(哈尔滨工业大学)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO

AI总结 TacMamba通过触觉历史压缩器和双阶段训练策略,实现高频率触觉数据与低频视觉推理的高效融合,提升实时任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00615 2026-03-03 cs.RO 79%

TGM-VLA: Task-Guided Mixup for Sampling-Efficient and Robust Robotic Manipulation

TGM-VLA:基于任务的混合学习用于高效且鲁棒的机器人操作

Fanqi Pu, Lei Jiang, Wenming Yang

机构 * Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(清华大学深圳国际研究生院) The National and Local Co-Build Humanoid Robotics Innovation Center(国家级与地方共建人形机器人创新中心)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO

AI总结 TGM-VLA通过优化关键帧采样策略和引入颜色反转投影模块,提升机器人操作任务的效率和鲁棒性。

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20577 2026-02-25 cs.CV 79%

Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion

高效的端到端自动驾驶:通过掩码视觉-语言-动作扩散

Jiaru Zhang, Manav Gagvani, Can Cui, Juntong Peng, Ruqi Zhang, Ziran Wang

机构 * Institute for Physical Artificial Intelligence (IPAI), Purdue University(物理人工智能研究所(IPAI)、普渡大学) College of Engineering, Purdue University(工程学院、普渡大学) Department of Computer Science, Purdue University(计算机科学系、普渡大学)

专题命中 VLA模型 :vision-language-action(title,abstract);分类 cs.CV

AI总结 MVLAD-AD通过掩码视觉-语言-动作扩散模型,提升自动驾驶的效率与规划精度,实现高效且可解释的端到端自动驾驶。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20383 2026-01-29 cs.CV 79%

HINT: Hierarchical Interaction Modeling for Autoregressive Multi-Human Motion Generation

HINT: 用于自回归多人体运动生成的分层交互建模

Mengge Liu, Yan Di, Gu Wang, Yun Qu, Dekai Zhu, Yanyan Li, Xiangyang Ji

机构 * Tsinghua University(清华大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

AI总结 HINT通过分层交互建模提出了一种自回归框架,实现多人体运动生成,提升了交互建模的精度和长序列的连贯性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19005 2026-01-28 cs.IR cs.LG stat.ML 79%

Recommending Composite Items Using Multi-Level Preference Information: A Joint Interaction Modeling Approach

利用多级偏好信息推荐复合物品:一种联合交互建模方法

Xuan Bi, Yaqiong Wang, Gediminas Adomavicius, Shawn Curley

机构 * Carlson School of Management, University of Minnesota(明尼苏达大学卡尔森管理学院) Leavey School of Business, Santa Clara University(圣克拉拉大学莱维商学院)

专题命中 VLA模型 :action model(title,abstract);分类 cs.LG

AI总结 本文提出JIMA方法,通过联合交互建模利用多级偏好信息,提升复合物品推荐的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11620 2025-12-15 cs.RO cs.SY eess.SY 79%

Architecting Large Action Models for Human-in-the-Loop Intelligent Robots

为具有人类在环的智能机器人构建大型动作模型

Kanisorn Sangchai, Methasit Boonpun, Withawin Kraipetchara, Paulo Garcia

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

AI总结 本文提出通过整合符号方法与现成模型构建可验证的神经符号智能机器人动作模型,以提升可靠性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12405 2025-11-18 cs.CV 79%

VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving

Hyunki Seong, Seongwoo Moon, Hojin Ahn, Jehun Kang, David Hyunchul Shim

机构 * School of Electrical Engineering, KAIST(韩国科学技术院电子工程学院)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.CV

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13397 2025-11-13 cs.CV 79%

Trustworthy Pedestrian Trajectory Prediction via Pattern-Aware Interaction Modeling

Kaiyuan Zhai, Juan Chen, Chao Wang, Zeyi Xu, Guoming Tang

机构 * Shanghai University(上海大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27383 2025-11-03 cs.AI 79%

Realistic pedestrian-driver interaction modelling using multi-agent RL with human perceptual-motor constraints

Yueyang Wang, Mehmet Dogar, Gustav Markkula

机构 * Institute for Transport Studies(交通研究所在) School of Computer Science(计算机科学学院)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14630 2025-09-19 cs.RO 79%

Toward Embodiment Equivariant Vision-Language-Action Policy

Anzhe Chen, Yifei Yang, Zhenjie Zhu, Kechun Xu, Zhongxiang Zhou, Rong Xiong, Yue Wang

专题命中 VLA模型 :vision-language-action(title,abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15885 2025-08-26 cs.AI 79%

VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making

Zuojin Tang, Bin Hu, Chenyang Zhao, De Ma, Gang Pan, Bin Liu

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Home Robotics Lab, E-surfing Digital Life Technology Co., Ltd., China Telecom(E-surfing数字生活技术有限公司) Zhejiang Lab(浙江实验室) Trinity College Dublin(都柏林大学)

专题命中 VLA模型 :action model(title);VLA(abstract);分类 cs.AI

Comments Accepted by EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00920 2025-08-13 physics.chem-ph cs.LG 79%

Uni-Mol3: A Multi-Molecular Foundation Model for Advancing Organic Reaction Modeling

Lirong Wu, Junjie Wang, Zhifeng Gao, Xiaohong Ji, Rong Zhu, Xinyu Li, Linfeng Zhang, Guolin Ke, Weinan E

机构 * AI for Science Institute(人工智能科学研究院) DP Technology(DP技术) College of Chemistry and Molecular Engineering(化学与分子工程学院) School of Mathematical Sciences(数学科学学院) Center for Machine Learning Research(机器学习研究中心)

专题命中 VLA模型 :action model(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18775 2025-07-28 cs.AI cs.SE 79%

Initial Steps in Integrating Large Reasoning and Action Models for Service Composition

Ilche Georgievski, Marco Aiello

机构 * University of Stuttgart(斯图加特大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

Comments 16 pages, 3 figures, 19th Symposium and Summer School on Service-Oriented Computing (SummerSOC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23121 2025-04-01 cs.CV 79%

Efficient Explicit Joint-level Interaction Modeling with Mamba for Text-guided HOI Generation

Guohong Huang, Ling-An Zeng, Zexin Zheng, Shengbo Gu, Wei-Shi Zheng

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

Comments Accepted to ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13006 2025-02-19 cs.AI 79%

Integrating Reinforcement Learning, Action Model Learning, and Numeric Planning for Tackling Complex Tasks

Yarin Benyamin, Argaman Mordoch, Shahaf S. Shperberg, Roni Stern

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10047 2025-01-14 cs.AI 79%

Large Action Models: From Inception to Implementation

Lu Wang, Fangkai Yang, Chaoyun Zhang, Junting Lu, Jiaxu Qian, Shilin He, Pu Zhao, Bo Qiao, Ray Huang, Si Qin, Qisheng Su, Jiayi Ye, Yudi Zhang, Jian-Guang Lou, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

Comments 25pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09631 2024-12-04 cs.AI 79%

Action Model Learning with Guarantees

Diego Aineto, Enrico Scala

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

Journal ref Proceedings of the International Conference on Principles of Knowledge Representation and Reasoning, 21(1), 801-811, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03307 2024-11-25 cs.RO cs.SY eess.SY 79%

Bi-level Trajectory Optimization on Uneven Terrains with Differentiable Wheel-Terrain Interaction Model

Amith Manoharan, Aditya Sharma, Himani Belsare, Kaustab Pal, K. Madhava Krishna, Arun Kumar Singh

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

Comments 8 pages, 7 figures, submitted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08848 2024-11-19 cs.RO 79%

Learning Spatial Bimanual Action Models Based on Affordance Regions and Human Demonstrations

Björn S. Plonka, Christian Dreher, Andre Meixner, Rainer Kartmann, Tamim Asfour

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

Comments 8 pages, accepted for publication at Humanoids 2024 - Copyright IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23218 2024-10-31 cs.CL cs.CV cs.HC 79%

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, Yu Qiao

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21342 2024-10-30 cs.MA cs.AI 79%

Heterogeneous Interaction Modeling With Reduced Accumulated Error for Multi-Agent Trajectory Prediction

Siyuan Chen, Jiahai Wang

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

Comments 20 pages, accepted by IEEE TNNLS

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07875 2024-10-10 cs.RO cs.AI cs.CL cs.CV cs.LG 79%

Generative Image as Action Models

Mohit Shridhar, Yat Long Lo, Stephen James

专题命中 VLA模型 :action model(title);分类 cs.RO、cs.CV、cs.AI

Comments CoRL 2024. Website, code, checkpoints: https://genima-robot.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17366 2024-09-10 cs.CV 79%

Generative Hierarchical Temporal Transformer for Hand Pose and Action Modeling

Yilin Wen, Hao Pan, Takehiko Ohkawa, Lei Yang, Jia Pan, Yoichi Sato, Taku Komura, Wenping Wang

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

Comments Accepted by ECCV HANDS Workshop 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09069 2024-07-19 cs.CV 79%

Dyadic Interaction Modeling for Social Behavior Generation

Minh Tran, Di Chang, Maksim Siniukov, Mohammad Soleymani

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

Comments The first two authors contribute equally. The paper is accepted by ECCV 2024. Project Page: https://boese0601.github.io/dim/ Code: https://github.com/Boese0601/Dyadic-Interaction-Modeling

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07067 2024-06-12 cs.IR cs.AI 79%

TIM: Temporal Interaction Model in Notification System

Huxiao Ji, Haitao Yang, Linchuan Li, Shunyu Zhang, Cunyi Zhang, Xuanping Li, Wenwu Ou

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08629 2024-05-27 cs.CV 79%

Scaling Up Dynamic Human-Scene Interaction Modeling

Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, Siyuan Huang

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14645 2024-03-25 cs.CY cs.AI 79%

Designing Multi-Step Action Models for Enterprise AI Adoption

Shreyash Mishra, Shrey Shah, Rex Pereira

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09957 2024-03-13 cs.CV 79%

Self-paced Multi-grained Cross-modal Interaction Modeling for Referring Expression Comprehension

Peihan Miao, Wei Su, Gaoang Wang, Xuewei Li, Xi Li

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV

Comments Accepted by TIP

详情

展开后加载摘要…

URL PDF HTML 收藏