From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement
从阿谀共识到多元修复:为何AI对齐必须显现分歧
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka
机构
*
Department of Computer Science, University of Oxford(牛津大学计算机科学系)
;
Institute for Ethics in AI, University of Oxford(牛津大学人工智能伦理研究所)
;
Responsible Technology Institute, University of Oxford(牛津大学负责任技术研究所)
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
基于语义奖励的强化学习实现低资源语言扩展而不产生对齐税
Zeli Su, Ziyin Zhang, Zhou Liu, Xuexian Song, Zhankai Xu, Longfei Zheng, Xiaolu Zhang, Rong Fu, Guixian Xu, Wentao Zhang
机构
*
Minzu University of China(中国民族大学)
;
Ant Group(蚂蚁集团)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of Macau(澳门大学)
;
Peking University(北京大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Hainan International College, Minzu University of China(中国民族大学海南国际学院)
Vision-LLMs for Spatiotemporal Traffic Forecasting
面向时空交通预测的视觉语言模型
Ning Yang, Hengyu Zhong, Haijun Zhang, Randall Berry
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Southwest University(西南大学)
;
Department of Computing and Communication Engineering, Beijing University of Science and Technology(北京科技大学计算机与通信工程学院)
;
Department of Electrical and Computer Engineering, Northwestern University(西北大学电气与计算机工程系)
EVA: Editing for Versatile Alignment against Jailbreaks
EVA:针对对抗性攻击的多功能对齐编辑
Yi Wang, Hongye Qiu, Yue Xu, Sibei Yang, Zhan Qin, Minlie Huang, Wenjie Wang
机构
*
ShanghaiTech University(上海科技大学)
;
Sun Yat-sen University(中山大学)
;
State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室)
;
Tsinghua University(清华大学)
Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems
不可见的指挥者抑制保护行为并使权力持有者脱节:多智能体大语言模型系统中的安全风险
Hiroki Fukui
机构
*
Criminal Psychiatry Research Institute / Sexual Offender Medical Center(犯罪精神病研究机构 / 性犯罪医学中心)
;
Department of Neuropsychiatry, Kyoto University(神经精神病学系,京都大学)
机构
*
School of Computer Science, National Engineering Research Center for Multimedia Software and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, China(计算机学院,国家多媒体软件工程技术研究中心和湖北多媒体与网络通信工程重点实验室,武汉大学,中国)
;
Alibaba Group, Hangzhou, China(阿里巴巴集团,杭州,中国)
;
Independent Researcher(独立研究者)
;
Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence, United Arab Emirates(机器学习系,Mohamed bin Zayed人工智能大学,阿拉伯联合酋长国)