arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9800 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9132 篇

2605.23987 2026-05-26 cs.AI cs.RO 81%

Beyond Predefined Learning Objects: A Thinking-Learning Interaction Model for Up-to-Date Autonomous Robot Learning

超越预定义学习对象:面向最新自主机器人学习的思维-学习交互模型

Hong Su

机构 * School of Computer Science, Chengdu University of Information Technology(成都信息科技大学计算机学院)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

AI总结 针对自主机器人在开放环境中无法依赖预定义学习对象的问题,提出一种思维-学习交互模型,通过思维指导学习(识别变化、选择证据、组织训练、规划验证)和学习促进思维(更新知识、经验、策略、推理)的双向机制,实现输入特征发现、输出类别扩展、模型更新和动作例程重构,实验验证了模型在特征适应、新类别形成、模型更新和动作优化上的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19352 2026-05-20 q-bio.NC cs.AI cs.LG 81%

Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay

在自然主义游戏过程中,视觉语言和动作模型的推理与动作表示的脑部对齐

Subba Reddy Oota, Anant Khandelwal, Khushbu Pahwa, Satya Sai Srinath Namburi, Tanmoy Chakraborty, Bapi S. Raju, Manish Gupta

机构 * Independent(独立) Microsoft Research(微软研究院) AWS AI Labs(AWS人工智能实验室) GE HealthCare(通用电气医疗) IIT Delhi(德里理工学院) IIIT-Hyderabad(海得拉巴理工学院) Microsoft(微软)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了在自然主义游戏过程中,视觉语言模型和大动作模型的推理与动作表示在脑部活动中的对齐情况,发现动作聚焦和推理聚焦的提示影响模型内部表示与fMRI脑活动的对齐程度。

Comments 21 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22963 2026-05-12 cs.RO cs.AI 81%

Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary

通过自由形式语言控制人形机器人:一个统一动作词汇的大型语言动作模型

Zhirui Liu, Kaiyang Ji, Ke Yang, Yahao Fan, Jingyi Yu, Ye Shi, Jingya Wang

机构 * ShanghaiTech University(上海科技大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出Humanoid-LLA模型,通过统一的人形机器人动作词汇解决语言与动作数据稀缺及物理稳定性问题,实现自由语言指令到全身动作的直接转换,提升人形机器人在真实环境中的泛化能力和动作多样性。

Comments Project page: https://humanoidlla.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06222 2026-05-12 cs.RO cs.AI 81%

When to Trust Imagination: Adaptive Action Execution for World Action Models

何时相信想象:为世界动作模型的自适应动作执行

Rui Wang, Yue Zhang, Jiehong Lin, Kuncheng Luo, Jianan Wang, Zhongrui Wang, Xiaojuan Qi

机构 * Southern University of Science and Technology(南方科技大学) The University of Hong Kong(香港大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出FFDC和混合时间 horizon训练,通过预测-观察一致性实现自适应动作执行,提升机器人在复杂场景中的鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07514 2026-05-11 cs.RO cs.CV 81%

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models

未来是否兼容?世界动作模型中的动态一致性诊断

Bo-Kai Ruan, Teng-Fang Hsiao, Ling Lo, Hong-Han Shuai

机构 * National Yang Ming Chiao Tung University

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.CV

AI总结 本文研究了世界动作模型中动作-状态一致性作为可靠性指标的重要性,提出了一种无需额外训练的共识策略提升规划效果。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11135 2026-04-14 cs.RO cs.LG 81%

AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps

AIM: 基于意图的统一世界动作建模与空间价值图

Liaoyuan Fan, Zetian Xu, Chen Cao, Wenyao Zhang, Mingqi Yuan, Jiayu Chen

机构 * INFIFORCE Intelligent Technology Co., Ltd.(英飞睿智科技有限公司) The University of Hong Kong(香港大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.LG

AI总结 AIM通过显式空间接口解决视频模型与动作生成间的结构不匹配问题,利用预训练视频生成模型和意图因果注意力机制,实现高成功率的机器人控制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25661 2026-04-08 cs.RO cs.CV 81%

Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance

Fast-dVLA:加速离散扩散VLA以实现实时性能

Wenxuan Song, Jiayi Chen, Shuai Chen, Jingbo Wang, Pengxiang Ding, Han Zhao, Yikai Qin, Xinhu Zheng, Donglin Wang, Yan Wang, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ShanghaiTech University(上海科技大学) Shanghai Institute of Technical Physics, CAS(中国科学院上海技术物理研究所) AIR, Tsinghua University(清华大学智能产业研究院) Westlake University(西湖大学) Zhejiang University(浙江大学)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.CV

AI总结 本文提出一种方法,通过分离辅助任务训练目标,在参数空间中提升通用能力与任务特定动作分布,从而在减少计算开销的同时提高性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16195 2026-03-19 cs.CV cs.RO 81%

S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight

S-VAM:通过自蒸馏几何和语义前瞻性构建快捷视频动作模型

Haodong Yan, Zhide Zhong, Jiaguan Zhu, Junjie He, Weilin Yuan, Wenxuan Song, Xin Gong, Yingjie Cai, Guanyi Zhao, Xu Yan, Bingbing Liu, Ying-Cong Chen, Haoang Li

机构 * The Hong Kong University of Science Technology (Guangzhou) Huawei Foundation Model Department

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.CV

AI总结 S-VAM通过自蒸馏策略实现单次前向传递生成几何和语义表示,提升动作预测效率,实验证明在复杂环境中优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13335 2026-03-17 cs.CV cs.AI 81%

Information-Theoretic Constraints for Continual Vision-Language-Action Alignment

信息论视角下的持续视觉-语言-动作对齐约束

Libang Zhao, Qixin Zeng, Hongyin Zhang, Donglin Wang

机构 * Westlake University, China(西湖大学) University of Southampton, UK(南安普顿大学)

专题命中 VLA模型 :vision-language-action(title);VLA(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Info-VLA框架,通过两个互补约束保持跨模态信息结构,解决持续学习中视觉语言动作对齐退化问题,实验表明其在任务保持与适应性上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15882 2026-02-19 cs.RO cs.AI 81%

FUTURE-VLA: Forecasting Unified Trajectories Under Real-time Execution

FUTURE-VLA:在实时执行下统一轨迹预测

Jingjing Fan, Yushan Liu, Shoujie Li, Botao Ren, Siyuan Li, Xiao-Ping Zhang, Wenbo Ding, Zhidong Deng

机构 * Department of Automation, Tsinghua University(自动化系,清华大学) Shenzhen International Graduation School, Tsinghua University(深圳国际研究生院,清华大学) Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.AI

AI总结 FUTURE-VLA通过实时执行实现统一轨迹预测,采用双效效率范式提升时空信息密度,实现实时预测与人机交互机制,取得多项任务的高成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00564 2026-01-30 cs.LG cs.AI 81%

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

通过联合优化的世界-动作模型扩展离线模型基于的强化学习

Jie Cheng, Ruixi Qiao, Yingwei Ma, Binhua Li, Gang Xiong, Qinghai Miao, Yongbin Li, Yisheng Lv

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

AI总结 JOWA通过联合优化的世界-动作模型扩展离线RL,实现高效泛化和高性能

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16065 2026-01-23 cs.CV cs.RO 81%

DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models

DTP: 一种简单而有效的干扰令牌修剪框架用于视觉-语言动作模型

Chenyang Li, Jieyuan Liu, Bin Li, Bo Gao, Yilin Yuan, Yangfan He, Yuchen Li, Jingqun Tang

机构 * Australian National University(澳大利亚国立大学) University of California, San Diego(加州大学圣地亚哥分校) Chinese Academy of Sciences(中国科学院) Beijing Institute of Graphic Communication(北京印刷学院) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Baidu Search(百度搜索) Bytedance(字节跳动)

专题命中 VLA模型 :action model(title);VLA(abstract);分类 cs.RO、cs.CV

AI总结 DTP框架通过动态修剪干扰令牌提升视觉-语言动作模型的任务成功率,适用于多种新型VLA模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09749 2026-01-16 cs.SE cs.AI cs.LG 81%

R-LAM: Reproducibility-Constrained Large Action Models for Scientific Workflow Automation

R-LAM:具有可重复性约束的大型动作模型用于科学工作流自动化

Suriya Sureshkumar

机构 * 1, Department of AI \& Data Science, RMK Engineering College, Chennai, India

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

AI总结 R-LAM通过引入结构化动作模式、确定性执行策略和显式溯源跟踪,提升科学工作流自动化的可重复性和可靠性。

Comments 9 pages, 3 figures, 1 Table, 2 Artifacts

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08713 2025-12-05 cs.LG cs.AI cs.IR 81%

Beyond KAN: Introducing KarSein for Adaptive High-Order Feature Interaction Modeling in CTR Prediction

超越KAN:引入KarSein用于CTR预测中的自适应高阶特征交互建模

Yunxiao Shi, Wujiang Xu, Haimin Zhang, Qiang Wu, Min Xu

机构 * University of Technology Sydney(悉尼技术大学) Rutgers University(罗格斯大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

AI总结 KarSein通过自适应高阶特征交互建模方法,提升CTR预测的准确性和效率,同时保持参数紧凑与解释性。

Comments Under review by TOIS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19236 2025-11-25 cs.RO cs.AI 81%

SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control

SENTINEL:一种用于人形机器人全身控制的端到端语言-动作模型

Yuxuan Wang, Haobin Jiang, Shiqing Yao, Ziluo Ding, Zongqing Lu

机构 * Peking University(北京大学) BeingBeyond

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

AI总结 SENTINEL是一种端到端语言-动作模型,通过直接映射语言指令和本体感觉输入到低层动作,实现人形机器人全身控制,并支持多模态扩展。

Comments 23 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04055 2025-11-19 q-bio.QM cs.AI cs.LG 81%

Benchmark on Drug Target Interaction Modeling from a Drug Structure Perspective

Xinnan Zhang, Jialin Wu, Junyi Xie, Tianlong Chen, Kaixiong Zhou

机构 * University of Minnesota(明尼苏达大学) University of California San Diego(加州大学圣地亚哥分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校) North Carolina State University(北卡罗来纳州立大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15691 2025-11-13 cs.LG cs.AI 81%

What Do Latent Action Models Actually Learn?

Chuheng Zhang, Tim Pearce, Pushi Zhang, Kaixin Wang, Xiaoyu Chen, Wei Shen, Li Zhao, Jiang Bian

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments Accepted by NeurIPS-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18865 2025-09-24 cs.RO cs.LG 81%

Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation

Masato Kobayashi, Thanpimon Buamanee

机构 * D3 Center, The University of Osaka(大阪大学D3中心) Graduate School of Information Science and Technology, The University of Osaka(大阪大学信息科学与技术研究生院) Graduate School of Maritime Sciences, Kobe University(兵库大学海运研究生院)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14067 2025-09-22 cs.CV cs.AI 81%

VLA-Mark: A cross modal watermark for large vision-language alignment model

Shuliang Liu, Qi Zheng, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) University of Toronto(多伦多大学) Ant Group, Alibaba(蚂蚁集团,阿里巴巴) New York University Shanghai(纽约大学上海分校)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by the main conference, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16365 2025-09-15 cs.CV cs.AI 81%

JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Muyao Li, Zihao Wang, Kaichen He, Xiaojian Ma, Yitao Liang

机构 * Peking University(北京大学) BIGAI(大疆创新)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16806 2025-07-28 cs.RO cs.AI 81%

DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation

Jiangran Lyu, Ziming Li, Xuesong Shi, Chaoyi Xu, Yizhou Wang, He Wang

机构 * Center on Frontiers of Computing Studies, School of Computer Science, Peking University(前沿计算研究教育部重点实验室,计算机学院,北京大学) Galbot Inst. for Artificial Intelligence, Peking University(人工智能研究所,北京大学) State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

Comments Project Page:https://pku-epic.github.io/DyWA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10705 2025-07-16 cs.LG cs.AI 81%

Learning Safe Numeric Planning Action Models

Argaman Mordoch, Shahaf S. Shperberg, Roni Stern, Berndan Juba

机构 * Software and Information Systems Engineering, Ben-Gurion University of the Negev(本·古里安大学软件与信息系统工程系) Department of Computer Science and Engineering, Washington University(华盛顿大学计算机科学与工程系)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13152 2025-07-01 cs.CV cs.AI 81%

Interpretable Interaction Modeling for Trajectory Prediction via Agent Selection and Physical Coefficient

Shiji Huang, Lei Ye, Min Chen, Wenhai Luo, Dihong Wang, Chenqi Xu, Deyuan Liang

机构 * College of Computer Science and Technology, Zhejiang University of Technology(计算机科学与技术学院,浙江工业大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by International Conference on Intelligent Robots and Systems (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22391 2025-06-06 cs.LG cs.AI 81%

A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks

Thomas Schmied, Thomas Adler, Vihang Patil, Maximilian Beck, Korbinian Pöppel, Johannes Brandstetter, Günter Klambauer, Razvan Pascanu, Sepp Hochreiter

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02298 2025-06-04 cs.CL cs.AI cs.LG 81%

LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback

Thai Hoang, Kung-Hsiang Huang, Shirley Kokane, Jianguo Zhang, Zuxin Liu, Ming Zhu, Jake Grigsby, Tian Lan, Michael S Ryoo, Chien-Sheng Wu, Shelby Heinecke, Huan Wang, Silvio Savarese, Caiming Xiong, Juan Carlos Niebles

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments LAM Simulator framework for agentic data generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08613 2025-05-27 cs.CV cs.AI 81%

Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation

Zhe Dong, Yuzhe Sun, Tianzhu Liu, Wangmeng Zuo, Yanfeng Gu

机构 * School of Electronics and Information Engineering, Harbin Institute of Technology(电子信息工程学院,哈尔滨工业大学) Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Peng Cheng Laboratory, Shenzhen(鹏城实验室,深圳)

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12272 2025-05-20 cs.AI cs.LG 81%

Enhancing Knowledge Graph Completion with GNN Distillation and Probabilistic Interaction Modeling

Lingzhi Wang, Pengcheng Huang, Haotian Li, Yuliang Wei, Guodong Xin, Rui Zhang, Donglin Zhang, Zhenzhou Ji, Wei Wang

机构 * Shandong Key Laboratory of Industrial Network Security(山东工业网络安全重点实验室) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00200 2025-04-28 cs.RO cs.CV 81%

Unified Video Action Model

Shuang Li, Yihuai Gao, Dorsa Sadigh, Shuran Song

机构 * Stanford University(斯坦福大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.CV

Comments Project website: https://unified-video-action-model.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02572 2025-03-05 cs.RO cs.AI 81%

RaceVLA: VLA-based Racing Drone Navigation with Human-like Behaviour

Valerii Serpiva, Artem Lykov, Artyom Myshlyaev, Muhammad Haris Khan, Ali Alridha Abdulkarim, Oleg Sautenkov, Dzmitry Tsetserukou

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.AI

Comments 6 pages, 6 figures. Submitted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14795 2025-02-24 cs.RO cs.CV 81%

Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

Pengxiang Ding, Jianfei Ma, Xinyang Tong, Binghong Zou, Xinxin Luo, Yiguo Fan, Ting Wang, Hongchao Lu, Panzhong Mo, Jinxin Liu, Yuefan Wang, Huaicheng Zhou, Wenshuo Feng, Jiacheng Liu, Siteng Huang, Donglin Wang

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏