AlloEgo-VLM: Disambiguating Allocentric and Egocentric Reference Frames in Vision-Language Models
AlloEgo-VLM:在视觉语言模型中消除 allocentric( allocentric 即 allocentric 参考框架,又称 allocentric 坐标系,指以环境为中心的参考框架)与 egocentric( egocentric 即 egocentric 参考框架,又称自我中心坐标系,指以观察者自身为中心的参考框架)参考框架的歧义
机构 * National Yang Ming Chiao Tung University(国立阳明交通大学) ; Institute of Computer Science and Engineering(工程与计算机科学学院) ; College of Artificial Intelligence(人工智能学院)
专题命中 VLM训练与架构 :VLM(title,title_cn);vision-language model(title,abstract);分类 cs.CV
AI总结 针对VLMs理解空间语义时 allocentric 与 egocentric 参考框架的歧义问题,构建数据集AlloEgo-View并开发框架AlloEgo-VLM,经NVIDIA Isaac Sim平台验证其在具身机器人开放式物体搜索任务中的有效性。
Comments 28 pages, 9 figures. Project page and code available at this https URL (https://github.com/CKL9001/AlloEgo-VLM)