From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
从结构到协同:多模态大语言模型中视觉-语言感知范式演进综述
机构 * School of Computer Science, Sichuan University(四川大学计算机学院) ; School of Electronic and Computer Engineering, Peking University Shenzhen Graduate School(北京大学深圳研究生院电子与计算机工程学院) ; Institute of Artificial Intelligence (TeleAI), China Telecom and Northwestern Polytechnical University(中国电信与西北工业大学人工智能研究院(TeleAI))
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本文系统综述多模态大语言模型中统一视觉-语言感知的范式演进,提出五阶段分类法,梳理各阶段代表性方法,并指出开放挑战与未来方向。