ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
ProGAL-VLA: 通过前瞻性推理实现视觉-语言-动作模型中的 grounded 对齐
机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校)
专题命中 VLA模型 :VLA(title,abstract);vision-language-action(title);action model(title);vision language action(abstract)
AI总结 ProGAL-VLA 通过构建 3D 实体中心图、使用慢速规划器生成符号子目标,并通过 Grounding Alignment Contrastive 损失对齐实体,提升机器人在扰动下的鲁棒性,减少语言忽略并提高实体检索性能。