arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2026-03-10 至 2026-03-10 共收录 20 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 20 篇

2509.14117 2026-03-10 cs.RO 91%

GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model

GeoAware-VLA: 基于隐式几何的视觉-语言-动作模型

Ali Abouzeid, Malak Mansour, Qinbo Sun, Zezhou Sun, Dezhen Song

机构 * Department of Robotics, Mohamed bin Zayed University of Artificial Intelligence(机器人系,Mohamed bin Zayed人工智能大学)

专题命中 VLA模型 :VLA(title,abstract);vision-language-action(title,abstract);action model(title);分类 cs.RO

AI总结 GeoAware-VLA通过整合几何先验提升视觉-语言-动作模型的视角不变性,显著提高零样本泛化能力,并在现实机器人平台中取得显著效果。

Comments Under Review, Project Page https://alisharey.github.io/GeoAware-VLA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08361 2026-03-10 cs.CV 91%

$Δ$VLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation

$Δ$VLA:通过世界知识变化引导的视觉-语言-动作模型

Yijie Zhu, Jie He, Rui Shao, Kaishen Yuan, Tao Tan, Xiaochen Yuan, Zitong Yu

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Great Bay University(大湾大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Macao Polytechnic University(澳门理工学院)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);action model(title);分类 cs.CV

AI总结 $Δ$VLA通过建模世界知识变化,结合先验引导和潜在空间学习,提升机器人动作生成的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08113 2026-03-10 cs.CV 91%

SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving

SAMoE-VLA:一种面向自动驾驶的场景自适应混合专家视觉-语言-动作模型

Zihan You, Hongwei Liu, Chenxu Dang, Zhe Wang, Sining Ang, Aoqi Wang, Yan Wang

机构 * Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学) School of Instrument Science and Engineering, Southeast University(仪器科学与工程学院,东南大学) Zhili College, Tsinghua University(紫荆学院,清华大学) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) Department of Automation, University of Science and Technology of China(自动化学院,中国科学技术大学) Department of Automation, University of Science and Technology Beijing(自动化学院,北京科技大学)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);action model(title);分类 cs.CV

AI总结 SAMoE-VLA通过场景自适应混合专家机制提升自动驾驶中的视觉-语言-动作推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00919 2026-03-10 cs.RO 91%

Green-VLA: Staged Vision-Language-Action Model for Generalist Robots

Green-VLA: 通用机器人中的分阶段视觉-语言-动作模型

I. Apanasevich, M. Artemyev, R. Babakyan, P. Fedotova, D. Grankin, E. Kupryashin, A. Misailidi, D. Nerus, A. Nutalapati, G. Sidorov, I. Efremov, M. Gerasyov, D. Pikurov, Y. Senchenko, S. Davidenko, D. Kulikov, M. Sultankin, K. Askarbek, O. Shamanin, D. Statovoy, E. Zalyaev, I. Zorin, A. Letkin, E. Rusakov, A. Silchenko, V. Vorobyov, S. Sobolnikov, A. Postnikov

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);action model(title);分类 cs.RO

AI总结 Green-VLA是一种面向真实环境的分阶段视觉-语言-动作模型,通过多阶段训练和强化学习对齐,提升机器人在多样化形态上的泛化能力和性能。

Comments 22 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10932 2026-03-10 cs.CR cs.AI cs.RO 88%

DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models

DropVLA: 对视觉-语言-动作模型的行动级后门攻击

Zonghuan Xu, Jiayu Li, Yunhan Zhao, Xiang Zheng, Xingjun Ma, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身AI研究院,复旦大学) Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身AI重点实验室) City University of Hong Kong(香港城市大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO、cs.AI

AI总结 DropVLA是一种针对视觉-语言-动作模型的行动级后门攻击,通过有限的数据污染实现对特定动作的精准操控,验证了在现实环境中的有效性。

Comments 8 pages, 6 tables, 3 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21243 2026-03-10 cs.RO 88%

RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models

RetoVLA: 通过重用寄存器令牌在视觉-语言-动作模型中实现空间推理

Jiyeon Koo, Taewan Cho, Hyunjoon Kang, Eunseom Pyo, Tae Gyun Oh, Taeryang Kim, Andrew Jaeyong Choi

机构 * School of Computing, Gachon University(高丽大学计算机学院)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 RetoVLA通过重用注册令牌提升视觉-语言-动作模型的空间推理能力,实验显示在7自由度机械臂任务中平均成功率提升17.1%。

Journal ref 2026 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07404 2026-03-10 cs.RO cs.AI 86%

Adaptive Capacity Allocation for Vision Language Action Fine-tuning

面向视觉语言动作微调的自适应容量分配

Donghoon Kim, Minji Bae, Unghui Nam, Gyeonghun Kim, Suyun Lee, Kyuhong Shim, Byonghyo Shim

机构 * Department of Electrical and Computer Engineering, Seoul National University(电气与计算机工程系,首尔国立大学) Department of Computer Science and Engineering, Sungkyunkwan University(计算机科学与工程系,全北国立大学)

专题命中 VLA模型 :vision language action(title,abstract);VLA(abstract);action model(abstract);分类 cs.RO、cs.AI

AI总结 本文提出LoRA-SP方法,通过自适应秩微调提升视觉语言动作模型在多任务场景中的性能和泛化能力。

Comments ICRA 2026 (Official Code: https://github.com/dhkim-furiosa/LoRA-SP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03596 2026-03-10 cs.RO cs.LG 85%

MEM: Multi-Scale Embodied Memory for Vision Language Action Models

MEM:多尺度具身记忆用于视觉语言动作模型

Marcel Torne, Karl Pertsch, Homer Walke, Kyle Vedder, Suraj Nair, Brian Ichter, Allen Z. Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, Danny Driess

专题命中 VLA模型 :vision language action(title);action model(title);分类 cs.RO、cs.LG

AI总结 MEM通过结合多模态记忆实现长期视觉语言动作任务的高效执行与适应性策略调整。

Comments Website: https://pi.website/research/memory

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08122 2026-03-10 cs.RO 84%

Towards Human-Like Manipulation through RL-Augmented Teleoperation and Mixture-of-Dexterous-Experts VLA

通过强化学习增强的遥控操作与混合的灵巧专家VLA实现人机操作

Tutian Tang, Xingyu Ji, Wanli Xing, Ce Hao, Wenqiang Xu, Lin Shao, Cewu Lu, Qiaojun Yu, Jiangmiao Pang, Kaifeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Sharpa Shanghai AI Lab(上海人工智能实验室) Zhongguancun Academy(中关村学院) National University of Singapore(新加坡国立大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 VLA模型 :VLA(title,abstract);vision-language-action(abstract);分类 cs.RO

AI总结 本文提出IMCopilot和MoDE-VLA框架,通过强化学习和多模态融合提升机器人灵巧操作能力,实现接触丰富的手内任务性能提升。

Comments Project Homepage: https://sites.google.com/view/mode-vla

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07647 2026-03-10 cs.RO 83%

TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon Vision-Language-Action Manipulation

TempoFit: 无需训练的分层时间KV内存用于长时视觉-语言-动作操控

Jun Sun, Boyu Yang, Jiahao Zhang, Ning Ma, Chencheng Wu, Siqing Zhang, Yiou Huang, Qiufeng Wang, Shan Liang, Yaran Chen

机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(abstract);分类 cs.RO

AI总结 TempoFit通过无需训练的分层时间KV内存提升长时视觉-语言-动作操控性能,有效解决记忆缺失问题,提升任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07484 2026-03-10 cs.RO 83%

HSC-VLA: Hierarchical Scene-Clearing for Robust Bimanual Manipulation in Dense Clutter

HSC-VLA:分层场景清除用于密集杂乱环境中的鲁棒双臂操作

Zhen Liu, Xinyu Ning, Zhe Hu, XinXin Xie, Yitong Liu, Zhongzhu Pu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学)

专题命中 VLA模型 :VLA(title,abstract);action model(abstract);分类 cs.RO

AI总结 HSC-VLA 通过分层场景清除方法提升密集杂乱环境中的双臂操作鲁棒性,实现 86.7% 的高成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08124 2026-03-10 cs.RO cs.AI cs.LG 82%

SaiVLA-0: Cerebrum--Pons--Cerebellum Tripartite Architecture for Compute-Aware Vision-Language-Action

SaiVLA-0:基于大脑-脑桥-小脑三元架构的计算感知视觉-语言-动作

Xiang Shi, Wenlong Huang, Menglin Zou, Xinhai Sun

专题命中 VLA模型 :vision-language-action(title,abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 SaiVLA-0提出基于大脑-脑桥-小脑三元架构的计算感知视觉-语言-动作系统,通过模块化设计提升效率与稳定性,初步实验显示训练时间减少且成功率显著提升。

Comments 14 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08519 2026-03-10 cs.RO 70%

AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models

AtomVLA: 通过预测性潜在世界模型实现机器人操作的可扩展后训练

Xiaoquan Sun, Zetian Xu, Chen Cao, Zonghe Liu, Yihan Sun, Jingrui Pang, Ruijian Zhang, Zhen Yang, Kang Pang, Dingxin He, Mingqi Yuan, Jiayu Chen

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.RO

AI总结 AtomVLA通过预测性潜在世界模型实现机器人操作的可扩展后训练,提升长时间任务的鲁棒性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07093 2026-03-10 cs.CV 70%

Facial Expression Generation Aligned with Human Preference for Natural Dyadic Interaction

面向自然双人互动的人类偏好对齐的面部表情生成

Xu Chen, Rui Gao, Xinjie Zhang, Haoyu Zhang, Che Sun, Zhi Gao, Yuwei Wu, Yunde Jia

专题命中 VLA模型 :vision-language-action(abstract);action model(abstract);分类 cs.CV

AI总结 本文提出了一种基于人类反馈的面部表情生成方法,通过动作学习和强化学习策略实现自然双人互动中与人类偏好的对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09820 2026-03-10 cs.RO cs.AI 62%

ViLAM: Distilling Vision-Language Reasoning into Attention Maps for Social Robot Navigation

ViLAM:将视觉-语言推理转化为注意力图用于社交机器人导航

Mohamed Elnoor, Kasun Weerakoon, Gershom Seneviratne, Jing Liang, Vignesh Rajagopal, Dinesh Manocha

机构 * Dept. of Electrical and Computer Engineering, University of Maryland, College Park, MD, USA(电气与计算机工程系,马里兰大学,College Park, MD, USA) Dept. of Computer Science, University of Maryland, College Park, MD, USA(计算机科学系,马里兰大学,College Park, MD, USA) James Clark School of Engineering, University of Maryland, College Park, MD, USA(James Clark工程学院,马里兰大学,College Park, MD, USA)

专题命中 VLA模型 :action model(abstract);分类 cs.RO、cs.AI

AI总结 ViLAM通过蒸馏视觉-语言模型的注意力图,提升社交机器人导航的社会适应性与导航效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06623 2026-03-10 cs.LG 57%

Advances in GRPO for Generation Models: A Survey

生成模型中GRPO的进展:综述

Zexiang Liu, Xianglong He, Yangguang Li

机构 * SJTU(上海交通大学) THU(清华大学) CUHK(香港大学)

专题命中 VLA模型 :vision-language-action(abstract);分类 cs.LG

AI总结 本文综述了流GRPO在生成模型中的应用与进展,探讨了方法论改进和跨模态扩展,强调其作为通用对齐框架的重要性及未来挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06601 2026-03-10 cs.LG 57%

Switchable Activation Networks

可切换激活网络

Laha Ale, Ning Zhang, Scott A. King, Pingzhi Fan

机构 * Southwest Jiaotong University(西南交通大学) University of Windsor(温莎大学) Texas A&M University-Corpus Christi(德克萨斯大学阿姆斯特朗分校)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

AI总结 SWAN通过动态控制神经元激活,实现高效推理和紧凑模型部署,统一稀疏性、剪枝和自适应推理的优势。

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08335 2026-03-10 eess.SY cs.SY 50%

The coordination between TSO and DSO in the context of energy transition - A review

TSO和DSO在能源转型中的协调 - 一种综述

Hang Nguyen, Koen Kok, Trung Thai Tran, Phuong H. Nguyen

专题命中 VLA模型 :action model(abstract)

AI总结 本文综述了TSO和DSO在能源转型中协调方案的分类、评估及挑战,旨在探索有效利用灵活性资源以维持系统平衡的方法。

Comments Published in: 2024 59th International Universities Power Engineering Conference (UPEC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15478 2026-03-10 cs.GT 50%

Equal-Pay Contracts

等额支付合同

Michal Feldman, Yoav Gal-Tzur, Tomasz Ponitka, Maya Schlesinger

专题命中 VLA模型 :action model(abstract)

AI总结 本文研究等额支付合同,分析其在多智能体合同设计中的算法和难度结果,并解决两个开放问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06715 2026-03-10 q-bio.PE stat.CO 50%

Understanding and Managing Frogeye Leaf Spot through Network-Based Modeling in Soybean

通过豆类网络建模理解与管理蛙眼叶斑病

Chinthaka Weerarathna, Thien-Minh Le, Jin Wang

专题命中 VLA模型 :action model(abstract)

AI总结 本文提出基于网络的模型,用于理解并管理大豆蛙眼叶斑病,通过近似贝叶斯计算估计流行病学参数,发现早期有针对性的除草更有效。

Comments 22 pages, 7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏