Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images
视觉自我实现对齐:通过威胁相关图像塑造安全导向的人设
Qishun Yang, Shu Yang, Lijie Hu, Di Wang
机构
*
King Abdullah University of Science and Technology(国王阿卜杜勒·阿齐兹大学科学与技术学院)
;
Provable Responsible AI and Data Analytics Lab(可证责任AI与数据分析实验室)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
China University of Petroleum-Beijing at Karamay(北京石油大学克拉玛依校区)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
City University of Hong Kong(香港城市大学)
;
Baidu Inc(百度公司)
;
Minzu University of China(民族大学)
"Even GPT Can Reject Me": Conceptualizing Abrupt Refusal Secondary Harm (ARSH) and Reimagining Psychological AI Safety with Compassionate Completion Standard (CCS)
TanksWorld: A Multi-Agent Environment for AI Safety Research
Corban G. Rivera, Olivia Lyons, Arielle Summitt, Ayman Fatima, Ji Pak, William Shao, Robert Chalmers, Aryeh Englander, Edward W. Staley, I-Jeng Wang, Ashley J. Llorens
CommentsThe full journal version of this article (published in Proceedings of the ACM on Human-Computer Interaction 4, CSCW2) can be found at https://dl.acm.org/doi/10.1145/3415168. The article is public access
Commentsv2: substantially expanded and retitled. Adds unpublished results on the dynamic (self-modifying) case, deriving the persistence barrier from Rice's Theorem one level up; a supervisory-regress theorem linking the results to scalable oversight and Yampolskiy's verifier theory; and a unified treatment of all four barriers as one obstruction, the Expressivity Invariant
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
跨模态冲突下的全方位安全:漏洞、动态机制和高效对齐
Kun Wang, Zherui Li, Zhenhong Zhou, Yitong Zhang, Yan Mi, Kun Yang, Yiming Zhang, Junhao Dong, Zhongxiang Sun, Qiankun Li, Yang Liu
机构
*
Nanyang Technological University(南洋理工大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Tsinghua University(清华大学)
;
Fudan University(复旦大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Renmin University of China(中国人民大学)
What is Safety? Corporate Discourse, Power, and the Politics of Generative AI Safety
什么是安全?企业话语、权力与生成式人工智能安全的政治
Ankolika De, Gabriel Lima, Yixin Zou
机构
*
College of Information Sciences and Technology, The Pennsylvania State University(信息科学与技术学院,宾夕法尼亚州立大学)
;
Max Planck Institute for Security and Privacy(安全与隐私研究所)
Advancing LLM Safe Alignment with Safety Representation Ranking
Tianqi Du, Zeming Wei, Quan Chen, Chenheng Zhang, Yisen Wang
机构
*
State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(通用人工智能国家重点实验室,智能科学与技术学院,北京大学)
;
School of Mathematical Sciences, Peking University(数学科学学院,北京大学)
;
Institute for Artificial Intelligence, Peking University(人工智能研究院,北京大学)