Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
面向视觉语言模型的层次化通用多模态攻击
Peng-Fei Zhang, Zi Huang
机构
*
School of Electrical Engineering and Computer Science, the University of Queensland(电气工程与计算机科学学院,昆士兰大学)
;
Department of Computer Science, City University of Hong Kong(计算机科学系,香港城市大学)
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
ScalSelect: 可扩展的无训练多模态数据选择用于高效的视觉指令微调
Changti Wu, Jiahuai Mao, Yuzhuo Miao, Shijie Lian, Bin Yu, Xiaopeng Lin, Cong Huang, Lei Zhang, Kai Chen
机构
*
East China Normal University(华东师范大学)
;
Zhongguancun Academy(中关村学院)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Huazhong University of Science and Technology(华中科技大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)
XAI-CLIP: ROI-Guided Perturbation Framework for Explainable Medical Image Segmentation in Multimodal Vision-Language Models
XAI-CLIP: 通过区域感兴趣引导扰动框架实现多模态视觉-语言模型中可解释的医学图像分割
Thuraya Alzubaidi, Sana Ammar, Maryam Alsharqi, Islem Rekik, Muzammil Behzad
机构
*
King Fahd University of Petroleum and Minerals(国王法赫德石油和矿物大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Imperial College London(伦敦帝国学院)
;
KFUPM-SDAIA Joint Research Centre for Artificial Intelligence(KFUPM-SDAIA联合人工智能研究中心)
Latent Reconstruction from Generated Data for Multimodal Misinformation Detection
从生成数据中进行潜在重建用于多模态虚假信息检测
Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, Panagiotis C. Petrantonakis
机构
*
Information Technology Institute, Centre for Research & Technology, Hellas(信息科技研究所,研究中心,希腊)
;
Department of Electrical & Computer Engineering, Aristotle University of Thessaloniki(电气与计算机工程系,亚里士多德大学)
机构
*
Qing Yuan Research Institute, Shanghai Jiao Tong University(上海交通大学庆元研究院)
;
Shanghai Innovation Institute(上海创新研究院)
;
Department of Pathology, The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学生命科学与医学学院病理科)
;
Intelligent Pathology Institute, Division of Life Sciences and Medicine(生命科学与医学学院智能病理研究所)
;
Department of Pathology, Fudan University Shanghai Cancer Center(复旦大学上海癌症中心病理科)
;
Department of Oncology, Shanghai Medical College, Fudan University(复旦大学上海医学院肿瘤科)
;
Institute of Pathology, Fudan University(复旦大学病理研究所)
;
Department of Pathology, The First Affiliated Hospital with Nanjing Medical University(南京医科大学第一附属医院病理科)
机构
*
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(人工智能学院,南京航空航天大学)
;
Department of Orthopedics, Qilu Hospital, Shandong University(骨科部,齐鲁医院,山东大学)
;
Key Laboratory of Qingdao in Medicine and Engineering(医学与工程青岛重点实验室)
Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models
Junjie Chen, Xuyang Liu, Subin Huang, Linfeng Zhang, Hang Yu
机构
*
Anhui Polytechnic University (AHPU)(安徽工程大学)
;
Shanghai University (SHU)(上海大学)
;
Shanghai Jiao Tong University (SJTU)(上海交通大学)
;
Sichuan University (SCU)(四川大学)
机构
*
Brown University(布朗大学)
;
University of Bristol(布里斯托大学)
;
Columbia University(哥伦比亚大学)
;
University of Rochester(罗切斯特大学)
;
Rutgers University(罗格斯大学)
Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
Hao Tang, Shengfeng He, Jing Qin
机构
*
Centre for Smart Health, The Hong Kong Polytechnic University(香港理工大学智能健康中心)
;
School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算机与信息系统学院)