ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
ACoT-VLA:面向视觉-语言-动作模型的行动链式推理
机构 * Beihang University(北京航空航天大学) ; AgiBot
AI总结 本文提出ACoT-VLA,通过在动作空间中直接进行推理,改进视觉-语言-动作模型的行动生成,引入显式和隐式动作推理器,实验证明其在真实和仿真环境中的优越性。
Comments Accepted by Conference on Computer Vision and Pattern Recognition (CVPR) 2026