Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
分析扩散模型和自回归视觉语言模型在多模态嵌入空间中的表现
机构 * Nanyang Technological University(南洋理工大学) ; Yale University(耶鲁大学) ; NYU Shanghai(纽约大学上海分校) ; Alibaba-NTU Singapore Joint Research Institute(阿里-国立新加坡大学联合研究机构) ; University of the Chinese Academy of Sciences(中国科学院大学) ; Center for Data Science(数据科学中心) ; New York University(纽约大学)
专题命中 其他LLM :language model(title,abstract);large language model(abstract);foundation model(abstract);分类 cs.CL、cs.AI
AI总结 本文研究了多模态扩散模型在多模态嵌入任务中的表现,发现其在分类、VQA和检索任务中均逊于自回归VLM。