AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
AffordanceVLA:一种通过可供性感知理解赋能动作生成的视觉-语言-动作模型
机构 * Peking University(北京大学) ; Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; The Chinese University of Hong Kong(香港中文大学) ; Knowin AI
专题命中 VLA模型 :VLA(summary_cn,abstract);vision-language-action(title,abstract);action model(title);分类 cs.RO、cs.CV
AI总结 提出AffordanceVLA框架,通过引入结构化可供性预测作为任务导向的中间表示,解决VLA模型中语义空间与具身控制策略的结构不匹配问题,实现精确的感知-动作映射。
Comments Preprint. Code and project page are available. Code: https://github.com/Skywalker-yqz/AffordanceVLA Project page: https://skywalker-yqz.github.io/AffordanceVLA/