MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion
MGDT:具有关系自适应专家混合的MLLM引导扩散变压器用于多模态知识图谱补全
Xu Hou, Meiyu Liang, Wei Huang, Yawen Li, Zhe Xue, Wu Liu, Guanhua Ye, Lei Shi, Kangkang Lu
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Zhejiang University(浙江大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Communication University of China(中国传媒大学)
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
Zhengyao Lv, Tianlin Pan, Chenyang Si, Zhaoxi Chen, Wangmeng Zuo, Ziwei Liu, Kwan-Yee K. Wong
机构
*
The University of Hong Kong(香港大学)
;
Nanjing University(南京大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Nanyang Technological University(南洋理工大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
Comments14 pages, 4 figures, 13 tables. Code, evaluation harness, and the released Temporal LLLite adapter weights are at https://github.com/otanl/dreamlite-stream (also mirrored to Hugging Face and Zenodo)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
ICG: 通过基于MLLM的提示和个性化偏好对齐改进封面图像生成
Zhipeng Bian, Jieming Zhu, Qijiong Liu, Wang Lin, Guohao Cai, Zhaocheng Du, Jiacheng Sun, Zhou Zhao, Zhenhua Dong
机构
*
Huazhong University of Science and Technology(华中科技大学)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
Hong Kong Polytechnic University(香港理工大学)
;
Zhejiang University(浙江大学)
CommentsPublished in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 12268-12278, EMNLP 2025. Official version: https://doi.org/10.18653/v1/2025.emnlp-main.617
Journal refProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (Main Track) EMNLP 2025 12268-12278
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
Run Luo, Xiaobo Xia, Lu Wang, Longze Chen, Renke Shan, Jing Luo, Min Yang, Tat-Seng Chua
机构
*
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
NExT++ Research Center(NExT++研究中心)
;
National University of Singapore(新加坡国立大学)
专题命中
多模态生成
:any-to-any(title,abstract);multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract)
MIMIC: A Generative Multimodal Foundation Model for Biomolecules
MIMIC:一种生成式多模态基础模型用于生物分子
Siavash Golkar, Jake Kovalic, Irina Espejo Morales, Samuel Sledzieski, Minhuan Li, Ksenia Sokolova, Geraud Krawezik, Alberto Bietti, Claudia Skok Gibbs, Roman Klypa, Shengwei Xiong, Francois Lanusse, Liam Parker, Kyunghyun Cho, Miles Cranmer, Tom Hehir, Michael McCabe, Lucas Meyer, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Helen Qu, Jeff Shen, David Fouhey, Hadi Sotoudeh, Vikram Mulligan, Pilar Cossio, Sonya M. Hanson, Alisha N. Jones, Olga G. Troyanskaya, Shirley Ho
机构
*
Polymathic AI Center for Data Science, New York University(多学科人工智能数据科学中心,纽约大学)
;
Polymathic AI Department of Applied Physics, Yale University(多学科人工智能应用物理系,耶鲁大学)
;
Center for Computational Mathematics, Flatiron Institute(计算数学中心,Flatiron研究所)
;
Center for Computational Biology, Flatiron Institute(计算生物学中心,Flatiron研究所)
;
Princeton Precision Health, Princeton University(普林斯顿精准健康,普林斯顿大学)
;
Department of Chemistry, New York University(化学系,纽约大学)
;
Department of Computer Science, Princeton University(计算机科学系,普林斯顿大学)
;
Lewis-Sigler Institute for Integrative Genomics, Princeton University(刘易斯-西格尔整合基因组学研究所,普林斯顿大学)
;
Center for Computational Astrophysics, Flatiron Institute(计算天文学中心,Flatiron研究所)
;
Department of Astrophysical Sciences, Princeton University(天体物理科学系,普林斯顿大学)
;
Department of Physics, New York University(物理学系,纽约大学)
专题命中
多模态生成
:multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.AI