MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction
MERGE:面向人类机器人交互中多主体事件推理与 grounding 的引导视觉-语言模型
机构 * Honda Research Institute Europe(本田欧洲研究院) ; Honda Research Institute USA(本田美国研究院)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract)
AI总结 MERGE 通过引导视觉语言模型实现动态人类机器人交互中多主体事件的推理与 grounding,提升 situational awareness 与效率,基于 GROUND 数据集验证其性能提升。