DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
DeepThinkVLA: 提升视觉-语言-动作模型的推理能力
机构 * Department of Computer Science and Technology, Tsinghua University, China(清华大学计算机科学与技术系) ; Gaoling School of Artificial Intelligence, Renmin University of China, China(中国人民大学 Gallagher 人工智能学院) ; Beijing Zhongguancun Academy, China(北京中关村学院)
专题命中 VLA模型 :vision-language-action(title,abstract_cn);action model(title);VLA(abstract,abstract_cn);vision language action(abstract)
AI总结 本文通过系统实验发现,Co-T推理需满足解码对齐和因果对齐两个条件,提出DeepThinkVLA模型在LIBERO等任务中取得显著提升。
Comments 26 pages, 7 figures, conference