Using large language models for embodied planning introduces systematic safety risks
使用大型语言模型进行具身规划引入了系统性的安全风险
Tao Zhang, Kaixian Qu, Zhibin Li, Jiajun Wu, Marco Hutter, Manling Li, Fan Shi
机构
*
ETH Zurich(苏黎世联邦理工学院)
;
University College London(伦敦大学学院)
;
Stanford University(斯坦福大学)
;
Northwestern University(西北大学)
;
National University of Singapore(新加坡国立大学)
Zhanyu Liu, Qingguo Hu, Ante Wang, Chenqing Liu, Zhishang Xiang, Hui Li, Delai Qiu, Jinsong Su
机构
*
School of Informatics, Xiamen University(厦门大学信息学院)
;
Tsinghua University(清华大学)
;
Xiamen Unisound Intelligence Technology Co., Ltd(厦门Unisound智能科技有限公司)
;
Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室(厦门大学),中华人民共和国文化和旅游部,中国)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
面向大语言模型内在可解释性的探索:设计原则与架构的综述
Yutong Gao, Qinglin Meng, Yuan Zhou, Liangming Pan
机构
*
MOE Key Lab of Computational Linguistics, Peking University(计算语言学MOE实验室,北京大学)
;
Beijing Academy of Artificial Intelligence, Beijing, China(北京人工智能研究院,北京,中国)
;
Nanjing University of Science and Technology(南京理工大学)
;
Purdue University(普渡大学)
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
有害合规的不同路径:跨大语言模型劫持的行为主效应和机制差异
Md Rysul Kabir, Zoran Tiganj
机构
*
Department of Computer Science(计算机科学系)
;
Luddy School of Informatics, Computing, and Engineering(信息学、计算与工程学院)
;
Indiana University Bloomington(印第安纳大学布卢明顿分校)
机构
*
Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳))
;
City University of Macau, Macao SAR, China(澳门城市大学)
;
Peng Cheng Laboratory, Shenzhen, China(鹏城实验室)
Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding
请拒绝回答我!通过自适应对比解码缓解大语言模型中的过度拒绝问题
Yupeng Qi, Ziyu Lyu, Lixin Cui, Lu Bai, Feng Xia
机构
*
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学信息科学与技术学院深圳校区)
;
School of Information, Central University of Finance and Economics(中央财经大学信息学院)
;
School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
;
School of Computing Technologies, RMIT University(皇家墨尔本理工大学计算技术学院)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
揭示LLM安全对齐中的logit抑制漏洞
Yuxi Li, Yi Liu, Yuekang Li, Ling Shi, Gelei Deng, Shengquan Chen, Kailong Wang
机构
*
Huazhong University of Science and Technology(华中科技大学)
;
Nanyang Technological University(南洋理工大学)
;
University of New South Wales(新南威尔士大学)
;
Nankai University(南开大学)