SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
SaPaVe:迈向视觉-语言-动作模型中的主动感知与操作
机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(信息多媒体国家重点实验室,计算机科学学院,北京大学) ; School of Software, Beihang University(软件学院,北航) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院)
专题命中 VLA模型 :vision-language-action(title,abstract);action model(title,abstract);分类 cs.RO、cs.CV
AI总结 SaPaVe通过解耦相机与操作动作,结合自下而上训练策略,实现高效且可推广的主动感知与操作,在现实任务中成功率达31.25%。
Comments Accepted to CVPR 2026. See project page at https://lmzpai.github.io/SaPaVe