Beyond Referring Expressions: Scenario Comprehension Visual Grounding
超越指称表达:场景理解视觉 grounding
机构 * Rice University(莱斯大学) ; Johns Hopkins University(约翰霍普金斯大学) ; Northeastern University(东北大学)
AI总结 本文提出RSC基准,探索基于场景的视觉 grounding,通过角色、意图和关系上下文推断目标,而非显式命名。ScenGround方法结合监督预热与难度感知强化学习,提升模型在复杂场景下的表现。
Comments 20 pages, 18 figures, Project Page: https://catherine-r-he.github.io/RSC/