Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
抓取任意区域:迈向多模态大语言模型的精确、上下文像素理解
机构 * New Laboratory of Pattern Recognition (NLPR), State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences (CASIA)(模式识别新实验室、多模态人工智能系统国家重点实验室、自动化研究所、中国科学院(CASIA)) ; University of Chinese Academy of Sciences(中国科学院大学) ; Peking University(北京大学) ; Wuhan University(武汉大学) ; ByteDance(字节跳动)
AI总结 GAR通过RoI对齐特征重放技术实现多区域精准理解与复杂推理,提升多模态大语言模型的上下文感知能力。
Comments ICLR 2026 Camera Ready Version