DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
DeepThinkVLA: 提升视觉-语言-动作模型的推理能力
机构 * Department of Computer Science and Technology, Tsinghua University, China(清华大学计算机科学与技术系) ; Gaoling School of Artificial Intelligence, Renmin University of China, China(中国人民大学 Gallagher 人工智能学院) ; Beijing Zhongguancun Academy, China(北京中关村学院)
专题命中 推理与问题求解 :SFT(abstract,abstract_cn);分类 cs.AI、cs.LG
AI总结 本文通过系统实验发现,Co-T推理需满足解码对齐和因果对齐两个条件,提出DeepThinkVLA模型在LIBERO等任务中取得显著提升。
Comments 26 pages, 7 figures, conference