Massively Multimodal Foundation Models: A Framework for Capturing Interactions with Specialized Mixture-of-Experts
大规模多模态基础模型:一种捕捉交互的专用专家混合框架
机构 * Johns Hopkins University(约翰霍普金斯大学) ; University of Texas at Austin(德克萨斯大学奥斯汀分校) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title);cross-modal(abstract)
AI总结 本文提出了一种大规模多模态基础模型框架,通过量化模态间的时间依赖性,改进混合专家路由机制,提升跨模态交互处理能力。
Comments Published at International Conference on Learning Representations (ICLR) 2026 as a conference paper. 28 pages, 16 figures, 10 tables