Interpreting the linear structure of vision-language model embedding spaces
Isabel Papadimitriou, Huangyuan Su, Thomas Fel, Sham Kakade, Stephanie Gil
机构
*
Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University(哈佛大学自然与人工智能研究所)
;
Department of Computer Science, Harvard University(哈佛大学计算机科学系)
机构
*
College of Computer Science & Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
Transvascular Implantation Devices Research Institute and Liangzhu Laboratory(血管植入物研究机构和良渚实验室)
;
Ant Group(蚂蚁集团)
;
University of Notre Dame(圣母大学)
;
HKUST (Guangzhou)(香港科技大学(广州))
专题命中
VLM训练与架构
:MLLM(title);LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
University of Science and Technology Beijing(北京科技大学)
;
Hunan University of Science and Technology(湖南科技大学)
;
Shanxi University of Finance and Economics(山西财经大学)
;
Beijing Friendship Hospital, Capital Medical University(首都医科大学北京友谊医院)
Effortless Vision-Language Model Specialization in Histopathology without Annotation
Jingna Qiu, Nishanth Jain, Jonas Ammeling, Marc Aubreville, Katharina Breininger
机构
*
Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-厄林根-纽伦堡大学)
;
Ingolstadt University of Applied Sciences(因戈尔施塔特应用科学大学)
;
Flensburg University of Applied Sciences(弗拉森堡应用科学大学)
;
Julius-Maximilians-Universität Würzburg(朱利叶斯-马克斯-魏扎克大学)
Defining and Evaluating Visual Language Models' Basic Spatial Abilities: A Perspective from Psychometrics
Wenrui Xu, Dalin Lyu, Weihang Wang, Jie Feng, Chen Gao, Yong Li
机构
*
School of Architecture, Tsinghua University(清华大学建筑学院)
;
Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)
;
BNRist, Tsinghua University(清华大学BNRist)
专题命中
VLM训练与架构
:visual language model(title,abstract);VLM(abstract);分类 cs.CV
Journal refProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Volume 1: Long Papers (2025) 11571-11590
机构
*
Technical University of Munich, Germany(慕尼黑技术大学,德国)
;
Helmholtz Munich, Munich Center for Machine Learning, Germany(海德堡慕尼黑,慕尼黑机器学习中心,德国)
;
University of Tübingen, Tübingen AI Center, Germany(图宾根大学,图宾根人工智能中心,德国)
;
University of Trento, Italy(特伦托大学,意大利)
;
Beijing University of Posts and Telecommunications, China(北京邮电大学,中国)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);LLaVA(abstract);分类 cs.CV
Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter
Bo Jiang, Xueyang Ze, Beibei Wang, Xixi Wang, Xixi Wan, Bin Luo
机构
*
Anhui Provincial Key Laboratory of Multimodal Cognitive Computation(安徽省级多模态认知计算重点实验室)
;
School of Computer Science and Technology(计算机科学与技术学院)
;
Anhui University(安徽大学)
Bidirectional Prototype-Reward co-Evolution for Test-Time Adaptation of Vision-Language Models
Xiaozhen Qiao, Peng Huang, Jiakang Yuan, Xianda Guo, Bowen Ye, Chaocan Xue, Ye Zheng, Zhe Sun, Xuelong Li
机构
*
School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学)
;
Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信)
;
College of Future Information Technology, Fudan University(未来信息技术学院,复旦大学)
;
College of Computer Science, Wuhan University(计算机科学学院,武汉大学)
专题命中
VLM训练与架构
:vision-language model(title);vision language model(abstract);VLM(abstract);分类 cs.CV