Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
超越跨模态对齐:测量和利用模态差距在视觉-语言模型中
Hanqi Yan, Xiangxiang Cui, Lu Yin, Jindong Gu, Paul Pu Liang, Yulan He, Yifei Wang
机构
*
King’s College London(伦敦国王学院)
;
University of Surrey(萨里大学)
;
University of Oxford(牛津大学)
;
MIT CSAIL(麻省理工学院CSAIL实验室)
;
The Alan Turing Institute(阿兰·图灵研究院)
机构
*
Technical University of Munich, Germany(慕尼黑技术大学,德国)
;
Munich Center for Machine Learning, Germany(慕尼黑机器学习中心,德国)
;
Helmholtz Munich, Germany(海德堡-慕尼黑赫尔姆霍尔茨研究中心,德国)
;
University of Tübingen, Tübingen AI Center, Germany(图宾根大学,图宾根人工智能中心,德国)
;
University of Trento, Italy(特伦托大学,意大利)
;
Beijing University of Posts and Telecommunications, China(北京邮电大学,中国)
专题命中
VLM训练与架构
:vision language model(title,abstract)
Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model Adaptation
并非所有方向都重要:迈向结构化和任务感知的低秩模型适应
Xi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu, Zhuxuanzi Wang, Yanshu Li, Lin Zhao, Guosheng Hu, Tianyang Wang, Hao Xu
机构
*
University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)
;
University of Virginia(弗吉尼亚大学)
;
Tulane University(路易斯安那州立大学)
;
Yale University(耶鲁大学)
;
Brown University(布朗大学)
;
Northeastern University(东北大学)
;
University of Bristol(布里斯托尔大学)
;
Harvard University(哈佛大学)
专题命中
VLM训练与架构
:LLaVA(abstract,abstract_cn);vision language model(abstract);分类 cs.CV
Wentao Zhang, Qi Zhang, Mingkun Xu, Mu You, Henghua Shen, Zhongzhi He, Keyan Jin, Derek F. Wong, Tao Fang
机构
*
Business School, Shandong University of Technology(山东理工大学商学院)
;
Faculty of Data Science, City University of Macau(澳门城市大学数据科学学院)
;
Guangdong Institute of Intelligent Science and Technology(广东智能科学与技术研究院)
;
Macau Millennium College(澳门 millennium 学院)
;
Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系)
CommentsThis work is an expanded version of our prior paper published in the IEEE ICASSP 2026 conference arXiv:2512.24947, from 4 to 20+ pages, presenting a well-structured and principled framework, extensive experiments, and deeper insights. Tao Fang is the corresponding author
机构
*
Shandong Provincial Key Laboratory of New Power Distribution & Utilization Technology and Equipment, School of Electrical and Electronic Engineering, Shandong University of Technology(山东省新型电力分布与利用技术及设备重点实验室,山东理工大学电气与电子工程学院)
;
School of Computer Science and Technology, Shandong University of Technology(山东理工大学计算机科学与技术学院)
;
School of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院)
;
Pillar of Information Systems Technology and Design, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计 pillar)
;
School of Control Science and Engineering, Shandong University(山东大学控制科学与工程学院)
EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models
EmoBench-M:多模态大语言模型情感智能评估基准
He Hu, Lianzhong You, Hongbo Xu, Qianning Wang, Fei Richard Yu, Fei Ma, Zebang Cheng, Zheng Lian, Yucheng Zhou, Laizhong Cui
机构
*
Shenzhen University(深圳大学)
;
Guangdong Laboratory of Artificial Intelligence(广东人工智能与数字经济实验室)
;
Auckland University of Technology(奥克兰理工大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
SKL-IOTSC, CIS, University of Macau(澳门科学技术大学SKL-IOTSC、CIS)
专题命中
其他VLM
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval
AnalogRetriever: 为模拟电路检索学习跨模态表示
Yihan Wang, Lei Li, Yao Lai, Jing Wang, Yan Lu
机构
*
Tsinghua University(清华大学)
;
The University of Hong Kong(香港大学)
;
University of Cambridge(剑桥大学)
;
Nanjing University of Posts and Telecommunications(南京邮电大学)