VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
VLA-Thinker: 通过图像推理增强视觉-语言-动作模型
专题命中 机器人基础模型 :manipulation(abstract);robotic(abstract);分类 cs.RO、cs.AI、cs.CV
AI总结 本文提出VLA-Thinker框架,通过动态可调用的推理动作提升视觉语言动作模型的能力,通过两阶段训练流程提升操控性能,在LIBERO和RoboTwin 2.0基准上取得显著成果。
Comments We introduce VLA-Thinker, the first VLA model capable of thinking-with-image reasoning, which models visual perception as a dynamically invocable reasoning action, enabling Multimodal Embodied Chain-of-Thought