机构
*
School of Software Engineering, Dalian University(大连大学软件工程学院)
;
School of Mathematical Sciences, Dalian University of Technology(大连理工大学数学科学学院)
;
The Hong Kong Polytechnic University(香港理工大学)
;
University of California, San Francisco(加利福尼亚大学旧金山分校)
;
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP
Grad-ECLIP: 基于梯度的CLIP视觉与文本解释
Chenyang Zhao, Kun Wang, Janet H. Hsiao, Antoni B. Chan
机构
*
Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)
;
Division of Social Science and Department of Computer Science & Engineering, Hong Kong University of Science & Technology(香港科学与技术大学社会科学学院及计算机科学与工程系)
;
SenseTime Group Ltd(时光集团有限公司)
Journal refZhao C, Wang K, Hsiao J H, et al. Grad-eclip: Gradient-based visual and textual explanations for clip[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026
Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization
用于3D理解的参数高效CLIP适配:通过统一分词实现
Guofeng Mei, Qinfeng Xiao, Bin Ren, Luigi Riz, Juan Liu, Xiaoshui Huang, Xu Zheng, Nicu Sebe, Ming-Hsuan Yang, Fabio Poiesi
机构
*
Fondazione Bruno Kessler(布鲁诺·科塞拉基金会)
;
University of Trento(特伦托大学)
;
University of Pisa(比萨大学)
;
Beijing Forestry University(北京林业大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Hong Kong University of Science and Technology (GZ)(香港科技大学)
;
Shandong University(山东大学)
;
University of California, Merced(加州大学默塞德分校)
机构
*
Xiamen University(厦门大学)
;
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
China University of Mining and Technology(中国矿业大学)
;
Shenzhen Technology University(深圳技术大学)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.CV
Test-Time Adaptation via Cache Personalization for Facial Expression Recognition in Videos
基于缓存个性化的时间测试适应用于视频面部表情识别
Masoumeh Sharafi, Muhammad Osama Zeeshan, Soufiane Belharbi, Alessandro Lameiras Koerich, Marco Pedersoli, Eric Granger
机构
*
LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(LIVIA系统工程系,蒙特利尔工程学院,加拿大)
;
LIVIA, Dept. of Software and IT Engineering, ETS Montreal, Canada(LIVIA软件与信息工程系,蒙特利尔工程学院,加拿大)
Lorenz Hufe, Niclas Griesshaber, Gavin Greif, Sebastian Oliver Eck, Philip Torr
机构
*
Torr Vision Group, Department of Engineering Science, University of Oxford(牛津大学工程科学系托尔视觉组)
;
Oxford Centre for Economic and Social History, University of Oxford(牛津大学牛津经济与社会史中心)
;
Faculty of Music, University of Oxford(牛津大学音乐学院)
;
Fraunhofer HHI(弗劳恩霍夫海因里希·赫兹研究所)
CommentsAccepted to ICML 2026. This version updates the ICML submission with an optimized model checkpoint. Project page: https://omni-diffusion.github.io
Token Communications: A Large Model-Driven Framework for Cross-modal Context-aware Semantic Communications
令牌通信:一种用于跨模态上下文感知语义通信的大模型驱动框架
Li Qiao, Mahdi Boloursaz Mashhadi, Zhen Gao, Rahim Tafazolli, Mehdi Bennis, Dusit Niyato
机构
*
School of Information and Electronics, Beijing Institute of Technology, Beijing 100081, China(信息与电子学院,北京理工大学,北京)
;
GIC & 6GIC, Institute for Communication Systems (ICS), University of Surrey, Guildford, United Kingdom(5GIC与6GIC,通信系统研究所(ICS), Surrey大学, Guildford,英国)
;
Centre for Wireless Communications, University of Oulu, 90014 Oulu, Finland(无线通信中心,奥卢大学, Oulu,芬兰)
;
School of Computer Science and Engineering, Nanyang Technological University, Singapore 639798(计算机科学与工程学院,南洋理工大学, Singapore)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.CV