arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2026-03-09 至 2026-03-09 共收录 7 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 5 篇

2511.06202 2026-03-09 cs.RO 91%

ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval

ExpReS-VLA: 通过经验回放和检索专门化视觉-语言-动作模型

Shahram Najam Syed, Yatharth Ahuja, Arthur Jakobsson, Jeff Ichnowski

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);action model(title);分类 cs.RO

AI总结 ExpReS-VLA通过经验回放和检索增强,提升VLA模型在特定任务中的适应能力与性能。

Comments 8 pages, 4 figures, 3 tables, accepted to International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05982 2026-03-09 cs.RO cs.CV 84%

HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the Wild

HarvestFlex: 通过视觉-语言-动作策略适应实现草莓采摘

Ziyang Zhao, Shuheng Wang, Zhonghua Miao, Ya Xiong

机构 * The Intelligent Equipment Research Center, Beijing Academy of Agriculture and Forestry Sciences(北京农业与林业科学研究院智能装备研究所) The School of Mechanical Electrical Engineering and Automation, Shanghai University(上海大学机械电子工程与自动化学院)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(abstract);分类 cs.RO、cs.CV

AI总结 HarvestFlex通过视觉-语言-动作策略适应,在真实温室中实现了高效草莓采摘,成功率达74%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05815 2026-03-09 cs.RO 79%

Hierarchical Latent Action Model

层次化潜在动作模型

Hanjung Kim, Lerrel Pinto, Seon Joo Kim

机构 * Yonsei University(延世大学) New York University(纽约大学)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

AI总结 HiLAM通过建模长期时间信息,从无动作视频中发现高层次潜在技能,提升了动态技能发现的鲁棒性。

Comments ICLR 2026 Workshop - 2nd Workshop on World Models: Understanding, Modelling and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06058 2026-03-09 cs.RO 57%

RODEO: RObotic DEcentralized Organization

Milan Groshev, Eduardo Castelló Ferrer

机构 * CyPhy Life, Robotics & AI Lab, School of Science & Technology, IE University(CyPhy生命、机器人与人工智能实验室,科学与技术学院,IE大学)

专题命中 VLA模型 :action model(abstract);分类 cs.RO

Comments 8 pages, 6 figures, Accepted at IEEE International Conference on Robotics & Automation (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06392 2026-03-09 astro-ph.GA 50%

Molecular Clouds Resolved at the Onset of Cosmic Noon

宇宙午夜初期的分子云解析

Bjorn Emonts, Matthew Lehnert, Mingyu Li, Azia Robinson, Stephen Curran, Montserrat Villar-Martin, Chris Carilli, Raffaella Morganti, Ilsang Yoon, Pierre Guillard, George Miley, Reinout van Weeren, Zheng Cai

专题命中 VLA模型 :VLA(abstract)

AI总结 在宇宙午夜初期,研究人员在射电星系B2 0902+34中发现七个多分子云,通过光谱解析揭示其化学和物理特性,为研究早期宇宙恒星形成提供新视角。

Comments Accepted for publication in ApJ Letters (10 pages, 3 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 语言条件控制 1 篇

2509.14063 2026-03-09 cs.RO 57%

Language Conditioning Improves Accuracy of Aircraft Goal Prediction in Non-Towered Airspace

语言引导提升非塔台空域飞机目标预测的准确性

Sundhar Vinodh Sangeetha, Chih-Yuan Chiu, Sarah H. Q. Li, Shreyas Kousik

机构 * Georgia Institute of Technology(佐治亚理工学院) School of Aerospace Engineering(航空航天工程学院) School of Electrical and Computer Engineering(电气与计算机工程学院) School of Mechanical Engineering(机械工程学院)

专题命中 语言条件控制 :language-conditioned robot(abstract);分类 cs.RO

AI总结 本文提出利用语言信息提升非塔台空域飞机目标预测准确性的多模态框架,通过整合自然语言理解和空间推理,有效提高自主决策能力。

Comments The last two authors advised equally. Accepted to the 2026 IEEE International Conference on Robotics and Automation. 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 动作表示与策略 1 篇

2603.06049 2026-03-09 cs.CV cs.RO 81%

Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models

魔鬼在窄政策中:解锁驾驶VLA模型的探索

Canyu Chen, Yuguang Yang, Zhewen Tan, Yizhi Wang, Ruiyi Zhan, Haiyan Liu, Xuanyao Mao, Jason Bao, Xinyue Tang, Linlin Yang, Bingchuan Sun, Yan Wang, Baochang Zhang

机构 * National Superior College for Engineers, Beihang University(北京航空航天大学工程师学院) School of Electronic Information Engineering, Beihang University(北京航空航天大学电子信息工程学院) Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院) Lenovo Group Limited(联想集团有限公司) State Key Laboratory of Media Convergence and Communication, Communication University of China(中国传媒大学媒体融合与传播国家重点实验室)

专题命中 动作表示与策略 :VLA(title,abstract);分类 cs.RO、cs.CV

AI总结 Curious-VLA通过双阶段设计缓解探索-利用困境,提升驾驶VLA模型的探索潜力,实现Navsim基准的最优性能。

Comments Accepted by CVPR2026 findings

详情

展开后加载摘要…

URL PDF HTML 收藏