Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
Jinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang, Tat-Seng Chua, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
National University of Singapore(新加坡国立大学)
;
Southeast University(东南大学)
;
Wuhan AI Research(武汉人工智能研究所)
Mitigating Object Hallucination via Robust Local Perception Search
Zixian Gao, Chao Yang, Zhanhui Zhou, Xing Xu, Chaochao Lu
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Center for Future Media & School of Computer Science and Engineering, University of Electronic Science and Technology of China(未来媒体中心及电子科技大学计算机科学与工程学院)
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);分类 cs.CV
On the Importance of Text Preprocessing for Multimodal Representation Learning and Pathology Report Generation
Ruben T. Lucassen, Tijn van de Luijtgaarden, Sander P. J. Moonemans, Gerben E. Breimer, Willeke A. M. Blokx, Mitko Veta
机构
*
Dept. of Pathology, University Medical Center Utrecht(病理学系,乌得勒支大学医学中心)
;
Dept. of Biomedical Engineering, Eindhoven University of Technology(生物医学工程系,埃因霍温理工大学)
;
Dept. of Mathematics and Computer Science, Eindhoven University of Technology(数学与计算机科学系,埃因霍温理工大学)
AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders
Yuqi Zhang, Yuchun Miao, Zuchao Li, Liang Ding
机构
*
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)
;
The University of Sydney(悉尼大学)
机构
*
Global Institute of Future Technology(未来技术全球研究院)
;
Shanghai Jiao Tong University(上海交通大学)
;
College of Humanities(人文学院)
;
Department of Automation(自动化系)
;
School of Aeronautics and Astronautics(航空宇航学院)
Interpreting Social Bias in LVLMs via Information Flow Analysis and Multi-Round Dialogue Evaluation
Zhengyang Ji, Yifan Jia, Shang Gao, Yutao Yue
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Shandong University(山东大学)
;
Institute of Deep Perception Technology, JITRI(深度感知技术研究所)
专题命中
幻觉与鲁棒性
:vision language model(abstract);分类 cs.AI
Adversarial Attacks of Vision Tasks in the Past 10 Years: A Survey
Chiyu Zhang, Lu Zhou, Xiaogang Xu, Jiafei Wu, Zhe Liu
机构
*
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
The University of Hong Kong(香港大学)
;
Zhejiang Lab, Nanjing University of Aeronautics and Astronautics(浙江实验室,南京航空航天大学)
机构
*
Department of Computer Science, Virginia Tech(弗吉尼亚理工学院计算机科学系)
;
Virginia Tech Transportation Institute, Virginia Tech(弗吉尼亚理工学院交通研究所)
;
Department of Statistics, Virginia Tech(弗吉尼亚理工学院统计学系)