arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9119 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9119 篇

2603.16195 2026-03-19 cs.CV cs.RO 81%

S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight

S-VAM:通过自蒸馏几何和语义前瞻性构建快捷视频动作模型

Haodong Yan, Zhide Zhong, Jiaguan Zhu, Junjie He, Weilin Yuan, Wenxuan Song, Xin Gong, Yingjie Cai, Guanyi Zhao, Xu Yan, Bingbing Liu, Ying-Cong Chen, Haoang Li

机构 * The Hong Kong University of Science Technology (Guangzhou) Huawei Foundation Model Department

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.CV

AI总结 S-VAM通过自蒸馏策略实现单次前向传递生成几何和语义表示,提升动作预测效率,实验证明在复杂环境中优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13335 2026-03-17 cs.CV cs.AI 81%

Information-Theoretic Constraints for Continual Vision-Language-Action Alignment

信息论视角下的持续视觉-语言-动作对齐约束

Libang Zhao, Qixin Zeng, Hongyin Zhang, Donglin Wang

机构 * Westlake University, China(西湖大学) University of Southampton, UK(南安普顿大学)

专题命中 VLA模型 :vision-language-action(title);VLA(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Info-VLA框架,通过两个互补约束保持跨模态信息结构,解决持续学习中视觉语言动作对齐退化问题,实验表明其在任务保持与适应性上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15882 2026-02-19 cs.RO cs.AI 81%

FUTURE-VLA: Forecasting Unified Trajectories Under Real-time Execution

FUTURE-VLA:在实时执行下统一轨迹预测

Jingjing Fan, Yushan Liu, Shoujie Li, Botao Ren, Siyuan Li, Xiao-Ping Zhang, Wenbo Ding, Zhidong Deng

机构 * Department of Automation, Tsinghua University(自动化系,清华大学) Shenzhen International Graduation School, Tsinghua University(深圳国际研究生院,清华大学) Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.AI

AI总结 FUTURE-VLA通过实时执行实现统一轨迹预测,采用双效效率范式提升时空信息密度,实现实时预测与人机交互机制,取得多项任务的高成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00564 2026-01-30 cs.LG cs.AI 81%

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

通过联合优化的世界-动作模型扩展离线模型基于的强化学习

Jie Cheng, Ruixi Qiao, Yingwei Ma, Binhua Li, Gang Xiong, Qinghai Miao, Yongbin Li, Yisheng Lv

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

AI总结 JOWA通过联合优化的世界-动作模型扩展离线RL,实现高效泛化和高性能

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16065 2026-01-23 cs.CV cs.RO 81%

DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models

DTP: 一种简单而有效的干扰令牌修剪框架用于视觉-语言动作模型

Chenyang Li, Jieyuan Liu, Bin Li, Bo Gao, Yilin Yuan, Yangfan He, Yuchen Li, Jingqun Tang

机构 * Australian National University(澳大利亚国立大学) University of California, San Diego(加州大学圣地亚哥分校) Chinese Academy of Sciences(中国科学院) Beijing Institute of Graphic Communication(北京印刷学院) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Baidu Search(百度搜索) Bytedance(字节跳动)

专题命中 VLA模型 :action model(title);VLA(abstract);分类 cs.RO、cs.CV

AI总结 DTP框架通过动态修剪干扰令牌提升视觉-语言动作模型的任务成功率,适用于多种新型VLA模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09749 2026-01-16 cs.SE cs.AI cs.LG 81%

R-LAM: Reproducibility-Constrained Large Action Models for Scientific Workflow Automation

R-LAM:具有可重复性约束的大型动作模型用于科学工作流自动化

Suriya Sureshkumar

机构 * 1, Department of AI \& Data Science, RMK Engineering College, Chennai, India

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

AI总结 R-LAM通过引入结构化动作模式、确定性执行策略和显式溯源跟踪,提升科学工作流自动化的可重复性和可靠性。

Comments 9 pages, 3 figures, 1 Table, 2 Artifacts

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08713 2025-12-05 cs.LG cs.AI cs.IR 81%

Beyond KAN: Introducing KarSein for Adaptive High-Order Feature Interaction Modeling in CTR Prediction

超越KAN:引入KarSein用于CTR预测中的自适应高阶特征交互建模

Yunxiao Shi, Wujiang Xu, Haimin Zhang, Qiang Wu, Min Xu

机构 * University of Technology Sydney(悉尼技术大学) Rutgers University(罗格斯大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

AI总结 KarSein通过自适应高阶特征交互建模方法,提升CTR预测的准确性和效率,同时保持参数紧凑与解释性。

Comments Under review by TOIS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19236 2025-11-25 cs.RO cs.AI 81%

SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control

SENTINEL:一种用于人形机器人全身控制的端到端语言-动作模型

Yuxuan Wang, Haobin Jiang, Shiqing Yao, Ziluo Ding, Zongqing Lu

机构 * Peking University(北京大学) BeingBeyond

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

AI总结 SENTINEL是一种端到端语言-动作模型,通过直接映射语言指令和本体感觉输入到低层动作,实现人形机器人全身控制,并支持多模态扩展。

Comments 23 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04055 2025-11-19 q-bio.QM cs.AI cs.LG 81%

Benchmark on Drug Target Interaction Modeling from a Drug Structure Perspective

Xinnan Zhang, Jialin Wu, Junyi Xie, Tianlong Chen, Kaixiong Zhou

机构 * University of Minnesota(明尼苏达大学) University of California San Diego(加州大学圣地亚哥分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校) North Carolina State University(北卡罗来纳州立大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15691 2025-11-13 cs.LG cs.AI 81%

What Do Latent Action Models Actually Learn?

Chuheng Zhang, Tim Pearce, Pushi Zhang, Kaixin Wang, Xiaoyu Chen, Wei Shen, Li Zhao, Jiang Bian

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments Accepted by NeurIPS-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18865 2025-09-24 cs.RO cs.LG 81%

Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation

Masato Kobayashi, Thanpimon Buamanee

机构 * D3 Center, The University of Osaka(大阪大学D3中心) Graduate School of Information Science and Technology, The University of Osaka(大阪大学信息科学与技术研究生院) Graduate School of Maritime Sciences, Kobe University(兵库大学海运研究生院)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14067 2025-09-22 cs.CV cs.AI 81%

VLA-Mark: A cross modal watermark for large vision-language alignment model

Shuliang Liu, Qi Zheng, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) University of Toronto(多伦多大学) Ant Group, Alibaba(蚂蚁集团,阿里巴巴) New York University Shanghai(纽约大学上海分校)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by the main conference, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16365 2025-09-15 cs.CV cs.AI 81%

JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Muyao Li, Zihao Wang, Kaichen He, Xiaojian Ma, Yitao Liang

机构 * Peking University(北京大学) BIGAI(大疆创新)

专题命中 VLA模型 :VLA(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16806 2025-07-28 cs.RO cs.AI 81%

DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation

Jiangran Lyu, Ziming Li, Xuesong Shi, Chaoyi Xu, Yizhou Wang, He Wang

机构 * Center on Frontiers of Computing Studies, School of Computer Science, Peking University(前沿计算研究教育部重点实验室,计算机学院,北京大学) Galbot Inst. for Artificial Intelligence, Peking University(人工智能研究所,北京大学) State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

Comments Project Page:https://pku-epic.github.io/DyWA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10705 2025-07-16 cs.LG cs.AI 81%

Learning Safe Numeric Planning Action Models

Argaman Mordoch, Shahaf S. Shperberg, Roni Stern, Berndan Juba

机构 * Software and Information Systems Engineering, Ben-Gurion University of the Negev(本·古里安大学软件与信息系统工程系) Department of Computer Science and Engineering, Washington University(华盛顿大学计算机科学与工程系)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13152 2025-07-01 cs.CV cs.AI 81%

Interpretable Interaction Modeling for Trajectory Prediction via Agent Selection and Physical Coefficient

Shiji Huang, Lei Ye, Min Chen, Wenhai Luo, Dihong Wang, Chenqi Xu, Deyuan Liang

机构 * College of Computer Science and Technology, Zhejiang University of Technology(计算机科学与技术学院,浙江工业大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by International Conference on Intelligent Robots and Systems (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22391 2025-06-06 cs.LG cs.AI 81%

A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks

Thomas Schmied, Thomas Adler, Vihang Patil, Maximilian Beck, Korbinian Pöppel, Johannes Brandstetter, Günter Klambauer, Razvan Pascanu, Sepp Hochreiter

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02298 2025-06-04 cs.CL cs.AI cs.LG 81%

LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback

Thai Hoang, Kung-Hsiang Huang, Shirley Kokane, Jianguo Zhang, Zuxin Liu, Ming Zhu, Jake Grigsby, Tian Lan, Michael S Ryoo, Chien-Sheng Wu, Shelby Heinecke, Huan Wang, Silvio Savarese, Caiming Xiong, Juan Carlos Niebles

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments LAM Simulator framework for agentic data generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08613 2025-05-27 cs.CV cs.AI 81%

Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation

Zhe Dong, Yuzhe Sun, Tianzhu Liu, Wangmeng Zuo, Yanfeng Gu

机构 * School of Electronics and Information Engineering, Harbin Institute of Technology(电子信息工程学院,哈尔滨工业大学) Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Peng Cheng Laboratory, Shenzhen(鹏城实验室,深圳)

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12272 2025-05-20 cs.AI cs.LG 81%

Enhancing Knowledge Graph Completion with GNN Distillation and Probabilistic Interaction Modeling

Lingzhi Wang, Pengcheng Huang, Haotian Li, Yuliang Wei, Guodong Xin, Rui Zhang, Donglin Zhang, Zhenzhou Ji, Wei Wang

机构 * Shandong Key Laboratory of Industrial Network Security(山东工业网络安全重点实验室) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00200 2025-04-28 cs.RO cs.CV 81%

Unified Video Action Model

Shuang Li, Yihuai Gao, Dorsa Sadigh, Shuran Song

机构 * Stanford University(斯坦福大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.CV

Comments Project website: https://unified-video-action-model.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02572 2025-03-05 cs.RO cs.AI 81%

RaceVLA: VLA-based Racing Drone Navigation with Human-like Behaviour

Valerii Serpiva, Artem Lykov, Artyom Myshlyaev, Muhammad Haris Khan, Ali Alridha Abdulkarim, Oleg Sautenkov, Dzmitry Tsetserukou

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.AI

Comments 6 pages, 6 figures. Submitted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14795 2025-02-24 cs.RO cs.CV 81%

Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

Pengxiang Ding, Jianfei Ma, Xinyang Tong, Binghong Zou, Xinxin Luo, Yiguo Fan, Ting Wang, Hongchao Lu, Panzhong Mo, Jinxin Liu, Yuefan Wang, Huaicheng Zhou, Wenshuo Feng, Jiacheng Liu, Siteng Huang, Donglin Wang

专题命中 VLA模型 :VLA(title,abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14891 2025-02-13 cs.RO cs.CV 81%

Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation

Guokang Wang, Hang Li, Shuyuan Zhang, Di Guo, Yanhong Liu, Huaping Liu

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04408 2025-02-10 cs.LG cs.AI 81%

Transforming Multimodal Models into Action Models for Radiotherapy

Matteo Ferrante, Alessandra Carosi, Rolando Maria D Angelillo, Nicola Toschi

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06067 2024-10-10 cs.CV cs.LG 81%

Contrastive Learning to Fine-Tune Feature Extraction Models for the Visual Cortex

Alex Mulrooney, Austin J. Brockmeier

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03215 2024-09-06 cs.CL cs.AI cs.LG 81%

xLAM: A Family of Large Action Models to Empower AI Agent Systems

Jianguo Zhang, Tian Lan, Ming Zhu, Zuxin Liu, Thai Hoang, Shirley Kokane, Weiran Yao, Juntao Tan, Akshara Prabhakar, Haolin Chen, Zhiwei Liu, Yihao Feng, Tulika Awalgaonkar, Rithesh Murthy, Eric Hu, Zeyuan Chen, Ran Xu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Silvio Savarese, Caiming Xiong

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

Comments Technical report for the Salesforce xLAM model series

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13627 2024-07-16 cs.CV cs.AI 81%

Vamos: Versatile Action Models for Video Understanding

Shijie Wang, Qi Zhao, Minh Quan Do, Nakul Agarwal, Kwonjoon Lee, Chen Sun

专题命中 VLA模型 :action model(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ECCV 2024 (European Conference on Computer Vision). Code and models are released at https://brown-palm.github.io/Vamos/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17802 2024-05-29 cs.LG cs.AI q-bio.BM 81%

Multi-level Interaction Modeling for Protein Mutational Effect Prediction

Yuanle Mo, Xin Hong, Bowen Gao, Yinjun Jia, Yanyan Lan

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14055 2024-04-08 cs.RO cs.AI 81%

Transforming a Quadruped into a Guide Robot for the Visually Impaired: Formalizing Wayfinding, Interaction Modeling, and Safety Mechanism

J. Taery Kim, Wenhao Yu, Yash Kothari, Jie Tan, Greg Turk, Sehoon Ha

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI

Comments 16 pages, 8 figures

Journal ref Proceedings of The 7th Conference on Robot Learning, PMLR 229:2288-2303, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏