Gold Points Sniper: Self-guided Visual Reasoning in VLM for Fine-grained Action Understanding
金点狙击手:VLM中的自引导视觉推理用于细粒度动作理解
机构 * Beijing National Research Center for Information Science and Technology (BNRist), Department of Automation, Tsinghua University(清华大学自动化系北京信息科学与技术国家研究中心) ; State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院通用人工智能国家重点实验室)
专题命中 视觉推理 :VLM(title,title_cn);visual reasoning(title);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 提出金点狙击手框架,通过金点提取器、选择性苏格拉底提问器和语义蕴含评估器三个模块,增强轻量级VLM的细粒度人类动作理解能力,在CAP基准上达到接近GPT-4o的性能且事实准确性更优。