The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning
良性多语言微调的安全影响异质性
Will Hawkins, Kaivalya Rawal, Jonathan Rystrøm, Stratis Tsirtsis, Zihao Fu, Greta Warren, Ryan Brown, Eoin Delaney, Sandra Wachter, Brent Mittelstadt, Chris Russell
CommentsAccepted at the 43rd International Conference on Machine Learning (ICML 2026). 25 pages, 8 figures. Project page: https://crda-project.github.io
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
你只有一份工作:利用LLM隐藏表示进行每任务量化
Amit LeVi, Raz Lapid, Rom Himelstein, Chaim Baskin, Ravid Shwartz Ziv, Avi Mendelson
机构
*
Technion—Israel Institute of Technology(技术ion—以色列理工学院)
;
Zenity(Zenity公司)
;
DeepKeep(DeepKeep公司)
;
Ben-Gurion University of the Negev(巴伊兰大学)
;
New York University(纽约大学)
CommentsCode and data: https://github.com/gregfrank/how-alignment-routes. Accepted at the Mechanistic Interpretability Workshop at the 43rd International Conference on Machine Learning (ICML), 2026
DiscoGen: Procedural Generation of Algorithm Discovery Tasks in Machine Learning
DiscoGen:机器学习算法发现任务的程序生成
Alexander D. Goldie, Zilin Wang, Adrian Hayler, Deepak Nathani, Edan Toledo, Ken Thampiratwong, Aleksandra Kalisz, Michael Beukman, Hannah Erlebach, Alistair Letcher, Shashank Reddy, Clarisse Wibault, Theo Wolf, Charles O'Neill, Uljad Berdica, Nicholas Roberts, Saeed Rahmani, Roberta Raileanu, Shimon Whiteson, Jakob N. Foerster
机构
*
University of Oxford(牛津大学)
;
University of California, Santa Barbara(加州大学圣塔芭芭拉分校)
;
University College London(伦敦大学学院)
;
University of Wisconsin--Madison(威斯康星大学麦迪逊分校)
;
Delft University of Technology(代尔夫特理工大学)