Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
多模态大语言模型的演进安全态势:新兴威胁与防护措施综述
Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu, Yusen Zhang, Cheng Han, Yushun Dong, Jiaqi Wang
机构
*
University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)
;
NVIDIA(英伟达公司)
;
Penn State University(宾夕法尼亚州立大学)
;
Columbia University(哥伦比亚大学)
;
University of Missouri-Kansas City(密苏里大学堪萨斯分校)
;
Florida State University(佛罗里达州立大学)
;
Auburn University(奥本大学)
机构
*
Singapore University of Technology and Design(新加坡科技设计大学)
;
Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局)
;
Nanyang Technological University(南洋理工大学)
;
Chongqing University(重庆大学)
专题命中
幻觉与鲁棒性
:grounding(abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG
机构
*
Australian National University(澳大利亚国立大学)
;
The University Of Queensland(昆士兰大学)
;
Peking University(北京大学)
;
GE research(通用电气研究院)
;
CSIRO(澳大利亚联邦科学与工业研究组织)
机构
*
The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳))
;
School of Data Science, School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(数据科学学院、人工智能学院、香港中文大学(深圳))
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
当观察不足时:视觉注意结构揭示大语言模型中的幻觉
Fanpu Cao, Xin Zou, Xuming Hu, Hui Xiong
机构
*
Thrust of Artificial Intelligence, HKUST (Guangzhou)(人工智能前沿 thrust,香港科技大学(广州))
;
Department of Computer Science and Engineering, HKUST(计算机科学与工程系,香港科技大学)
专题命中
幻觉与鲁棒性
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
South China University of Technology(华南理工大学)
;
Institute for Super Robotics (Huangpu)(机器人研究所(黄埔))
;
Shanghai Jiao Tong University(上海交通大学)
;
Changsha University of Science and Technology(长沙理工大学)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
PromptGuard: 基于软提示的文本到图像模型不安全内容审查
Lingzhi Yuan, Xinfeng Li, Chejian Xu, Guanhong Tao, Xiaojun Jia, Yihao Huang, Wei Dong, Yang Liu, Bo Li
机构
*
Department of Computer Science, University of Maryland(计算机科学系,马里兰大学)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机科学与数据科学学院)
;
Kahlert School of Computing, The University of Utah(犹他大学Kahlert计算学院)
;
Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校Siebel计算与数据科学学院)
Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images
视觉自我实现对齐:通过威胁相关图像塑造安全导向的人设
Qishun Yang, Shu Yang, Lijie Hu, Di Wang
机构
*
King Abdullah University of Science and Technology(国王阿卜杜勒·阿齐兹大学科学与技术学院)
;
Provable Responsible AI and Data Analytics Lab(可证责任AI与数据分析实验室)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
China University of Petroleum-Beijing at Karamay(北京石油大学克拉玛依校区)
专题命中
幻觉与鲁棒性
:vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI