LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
LOCUS: 局部视觉线索搜索增强多模态大语言模型的细粒度感知
机构 * University of Science and Technology of China(中国科学技术大学) ; State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)
专题命中 其他多模态 :multimodal(title,abstract);MLLM(summary_cn);分类 cs.CV
AI总结 提出LOCUS训练框架,通过可验证的局部线索搜索代理任务,使MLLM内化细粒度证据选择,提升定位敏感视觉理解而不改变推理接口。