Comments17 pages, 14 figures, accepted to Computer Vision and Pattern Recognition Conference (CVPR) Workshops 2026. 5th MMFM Workshop: What is Next in Multimodal Foundation Models?
Journal refIn Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7415-7424) 2026
机构
*
Nepal Applied Mathematics and Informatics Institute for Research(尼泊尔应用数学与信息技术研究所)
;
GastroIntestinal Department, Dhulikhel Hospital(杜尔基hel医院消化内科)
;
Univesity of Lausanne(洛桑大学)
;
University of West Virginia(西弗吉尼亚大学)
;
University of Utah(犹他大学)
;
University of Aberdeen(阿伯丁大学)
SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials
SoM-1K:用于材料力学能力的千题基准数据集
Qixin Wan, Zilong Wang, Jingwen Zhou, Wanting Wang, Ziheng Geng, Jiachen Liu, Ran Cao, Lu Cheng
机构
*
College of Civil Engineering, Hunan University(湖南大学土木工程学院)
;
Department of Civil & Architectural Engineering, University of Miami(迈阿密大学土木与建筑工程系)
;
Department of Electrical and Computer Engineering, University of Miami(迈阿密大学电气与计算机工程系)
;
School of Architecture, University of Miami(迈阿密大学建筑学院)
;
Department of Computer Science, University of Illinois Chicago(伊利诺伊大学芝加哥分校计算机科学系)
专题命中
视觉推理
:vision language model(abstract,abstract_cn)
AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference
AdaVFM:通过LLM引导执行实现边缘智能的自适应视觉基础模型
Yiwei Zhao, Yi Zheng, Huapeng Su, Jieyu Lin, Stefano Ambrogio, Cijo Jose, Michael Ramamonjisoa, Patrick Labatut, Barbara De Salvo, Chiao Liu, Phillip B. Gibbons, Ziyun Li
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
Meta
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);分类 cs.CV、cs.LG
机构
*
National Key Laboratory of Big Data and Decision(大数据与决策国家重点实验室)
;
National University of Defense Technology(国防科技大学)
;
The Center for machine learning research(机器学习研究中心)
;
Peking University(北京大学)
;
Tsinghua University(清华大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
从表征互补性到双系统:协同VLM和纯视觉骨干网络用于端到端驾驶
Sining Ang, Yuguang Yang, Chenxu Dang, Canyu Chen, Cheng Chi, Haiyan Liu, Xuanyao Mao, Jason Bao, Xuliang, Bingchuan Sun, Yan Wang
机构
*
Department of Automation, University of Science and Technology of China(中国科学技术大学自动化系)
;
School of Electronic Information Engineering, Beihang University(北京航空航天大学电子信息工程学院)
;
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
;
National Superior College for Engineers, Beihang University(北京航空航天大学国家级工程师学院)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Lenovo Group Limited(联想集团有限公司)
;
Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院)