3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
3D CAVLA:利用深度和3D上下文来泛化视觉语言动作模型以应对未见任务
机构 * New York University Tandon School of Engineering(纽约大学坦登工程学院) ; New York University Courant Institute of Mathematical Sciences(纽约大学库朗数学科学研究所)
专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO
AI总结 本文提出3D-CAVLA框架,通过引入链式推理、深度感知和任务导向的感兴趣区域检测,提升视觉语言动作模型在未见任务中的泛化能力,实验表明其在模拟和现实任务中均表现出色。
Comments Accepted at the 1st Workshop on 3D LLM/VLA, CVPR 2025. This work has been submitted to the IEEE for possible publication