VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
VLA-Thinker: 通过图像推理增强视觉-语言-动作模型
专题命中 VLA模型 :VLA(title,abstract);vision-language-action(title,abstract);action model(title);分类 cs.RO、cs.CV、cs.AI
AI总结 本文提出VLA-Thinker框架,通过动态可调用的推理动作提升视觉语言动作模型的能力,通过两阶段训练流程提升操控性能,在LIBERO和RoboTwin 2.0基准上取得显著成果。
Comments We introduce VLA-Thinker, the first VLA model capable of thinking-with-image reasoning, which models visual perception as a dynamically invocable reasoning action, enabling Multimodal Embodied Chain-of-Thought