V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators
V-Reflection:将多模态大语言模型从被动观察者转变为主动提问者
机构 * AI Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)人工智能方向) ; International Digital Economy Academy(国际数字经济学院) ; MedVisAI Lab, Lee Kong Chian School of Medicine, Nanyang Technological University, and Centre of AI in Medicine(医学视觉人工智能实验室,南洋理工大学Lee Kong Chian医学院,以及人工智能在医学中的中心) ; South China University of Technology(华南理工大学) ; Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(香港理工大学电子与电气工程系)
专题命中 其他多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 本文提出V-Reflection框架,通过'思考后再观察'的视觉反思机制,使多模态大语言模型能主动提问视觉特征空间,提升细粒度任务的感知能力。
Comments Main paper 14 pages with supplementary 7 pages