TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models
TokenSwap:对大型视觉语言模型组合理解的后门攻击
Zhifang Zhang, Qiqi Tao, Jiaqi Lv, Na Zhao, Lei Feng, Joey Tianyi Zhou
机构
*
School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院)
;
School of Electrical Engineering and Computer Science, University of Queensland, Australia(昆士兰大学电气工程与计算机科学学院)
;
Design Pillar, Singapore University of Technology and Design(新加坡科技设计大学设计学院)
;
Centre for Frontier AI Research (CFAR), Agency for Science, Technology and Research (A*STAR), Singapore(前沿人工智能研究中心(CFAR))
;
Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR), Singapore(高性能计算研究所(IHPC))
机构
*
School of Information Science and Technology, Northwest University(西北大学信息科学与技术学院)
;
Laboratory of Intelligent Information Processing, Institute of Computing Technology(中国科学院计算技术研究所智能信息处理重点实验室)
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
RSICCLLM:用于遥感图像变化描述的多模态大语言模型
Yelin Wang, Zijia Song, Shuo Ye, Chuanguang Yang, Miaoyu Wang, Yong Xu, Zhulin An, Yongjun Xu, Zitong Yu
机构
*
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,北京,中国)
;
Great Bay University, Dongguan, China(东莞Great Bay大学,中国)
;
Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学深圳学院,中国)
;
Dongguan Key Laboratory for Intelligence and Information Technology, Dongguan, China(东莞智能与信息科技重点实验室,中国)
专题命中
VLM训练与架构
:multimodal large language model(title);vision-language model(abstract);分类 cs.CV
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
面向选择性多模态大语言模型遗忘的良性记忆遗忘
Zhen Zeng, Leijiang Gu, Zhangling Duan, Feng Li, Cees G. M. Snoek, Meng Wang, Zenglin Shi
机构
*
Hefei University of Technology(合肥工业大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
;
University of Amsterdam(阿姆斯特丹大学)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);分类 cs.AI
机构
*
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
Zhejiang Key Laboratory of Digital-Intelligence Service Technology(浙江省数字化服务技术重点实验室)
;
Hong Kong Polytechnic University(香港理工大学)
;
The University of Southern Queensland(南昆士兰大学)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);分类 cs.CV
Timage: A Generative Text-in-Image Paradigm for Fine-Tuning Vision-Language Models
Timage: 一种用于微调视觉语言模型的文本嵌入图像生成范式
Yifeng Wu, Huimin Huang, Ruiluo Wu, Chunyi Lin, Guanhua Chen, Xian Wu, Wang Song, Ruize Han
机构
*
Fudan University(复旦大学)
;
Shenzhen University of Advanced Technology(深圳先进技术大学)
;
Tencent Jarvis Lab(腾讯贾维斯实验室)
;
Southern University of Science and Technology(南方科技大学)
专题命中
VLM训练与架构
:vision-language model(title);multimodal large language model(abstract);分类 cs.CV
DREAM: Extending Vision-Language Models with Dual-Objective Encoding for Cross-Modal Retrieval
DREAM: 通过双目标编码扩展视觉-语言模型用于跨模态检索
Kaleem Ullah, Altaf Hussain, Muhammad Munsif, Sung Wook Baik
机构
*
Sejong University(世宗大学)
;
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
Ulsan National Institute of Science and Technology(乌山国立科学研究院)
Attention Alignment Between Humans and Vision-Language Models
人类与视觉语言模型之间的注意力对齐
Isaac R. Christian, Udith Haputhanthrige, Hanna Hornfeld, Declan Campbell, Samuel Nastase, Taylor Webb, Michael Graziano
机构
*
Princeton Neuroscience Institute, Princeton University(普林斯顿大学普林斯顿神经科学研究所)
;
Department of Psychology, Princeton University(普林斯顿大学心理学系)
;
Department of Computer Science, Princeton University(普林斯顿大学计算机科学系)
;
Department of Psychology and Center for Computational Language Sciences, University of Southern California(南加州大学心理学系与计算语言科学中心)
;
Department of Psychology, Université de Montréal(蒙特利尔大学心理学系)
CommentsMICCAI 2026 Early Accept; Project Page: https://tahakoleilat.github.io/Evi-Steer. This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution will be published as part of the MICCAI 2026 proceedings in October
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
FTibSuite:面向藏语视觉语言建模的综合资源套件
Guixian Xu, Yide Liang, Zeli Su, Xuexian Song, Ziyin Zhang, Yushuang Dong, Ting Zhang, Xu Han
机构
*
Hainan International College, Minzu University of China(民族大学海南国际学院)
;
School of Information Engineering, Minzu University of China(民族大学信息工程学院)
;
Shanghai Jiao Tong University(上海交通大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
Visual-Advantage On-Policy Distillation for Vision-Language Models
基于视觉优势的在线策略蒸馏用于视觉-语言模型
Ruiqi Liu, Xiaolei Lv, Gengsheng Li, Ximo Zhu, Zhiheng Wang, Zhengbo Zhang, Junkai Chen, Zhiheng Li, Bo Li, Jun Gao, Shu Wu
机构
*
Institute of Automation, CAS(中国科学院自动化研究所)
;
School of Advanced Interdisciplinary Sciences, UCAS(中国科学院大学(UCAS)先进交叉学科学院)
;
Hello Group Inc.(Hello集团有限公司)
;
Sun Yat-sen University(中山大学)
机构
*
School of Geosciences and Info-Physics, Central South University(地质科学与信息物理学院,中南大学)
;
School of Earth Sciences and Spatial Information Engineering, Hunan University of Science and Technology(地球科学与空间信息工程学院,湖南科技大学)
Comments4 pages , 3 figures , This paper has been submitted to the IEEE-affiliated ICBME Conference (Iran), 2025, and is currently under review. DOR number: [20.1001.2.0425023682.1404.10.1.440.7]
AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
AdaMMS: 为异构多模态大语言模型设计的模型融合方法
Yiyang Du, Xiaochen Wang, Chi Chen, Jiabo Ye, Yiru Wang, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Zhifang Sui, Maosong Sun, Yang Liu
机构
*
Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University(计算机科学与技术系,人工智能研究院,清华大学)
;
Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学)
;
State Key Laboratory of Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学)
;
School of Software Microelectronics, Peking University(软件微电子学院,北京大学)
;
Institute of Intelligent Computing, Alibaba Group(智能计算研究院,阿里巴巴集团)
;
Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国)
;
Jiangsu Collaborative Innovation Center for Language Competence, Jiangsu, China(江苏省语言能力协同创新中心,江苏,中国)
;
ModelTC Open Source Organization, Beijing, China(ModelTC开源组织,北京,中国)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);分类 cs.CV