InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
InternVL-U: 使统一多模态模型在理解、推理、生成和编辑方面更加普及
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; Fudan University(复旦大学) ; University of Science and Technology of China(中国科学技术大学) ; Shanghai Jiao Tong University(上海交通大学) ; South China University of Technology(华南理工大学) ; Nanjing University(南京大学) ; Xiamen University(厦门大学) ; CUHK MMLab(香港中文大学MMLab)
专题命中 VLM训练与架构 :InternVL(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
AI总结 InternVL-U通过轻量级统一多模态模型,在保持强大生成能力的同时,实现了在理解、推理、生成和编辑方面的高效平衡。
Comments technical report, 61 pages, https://github.com/OpenGVLab/InternVL-U