CommentsCode and data: https://github.com/gregfrank/how-alignment-routes. Accepted at the Mechanistic Interpretability Workshop at the 43rd International Conference on Machine Learning (ICML), 2026
The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning
良性多语言微调的安全影响异质性
Will Hawkins, Kaivalya Rawal, Jonathan Rystrøm, Stratis Tsirtsis, Zihao Fu, Greta Warren, Ryan Brown, Eoin Delaney, Sandra Wachter, Brent Mittelstadt, Chris Russell
Reinforcement Learning with a Bilevel World-Model Architecture for Scan-Order Optimisation in Laser Directed Energy Deposition
激光增材制造扫描顺序优化的强化学习:用于奖励和世界模型诊断的双层代理-有限元分析诊断框架
Xian Wu, Haoran Li, Yuanqi Chu, Dongbin Zhao, Bin Wang
机构
*
College of Engineering, Design and Physical Sciences, Brunel University London(布鲁内尔大学伦敦工程、设计与物理科学学院)
;
Pattern Recognition Laboratory, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别实验室)
;
ISIS Neutron and Muon Source, Science and Technology Facilities Council, Rutherford Appleton Laboratory(Rutherford Appleton实验室,科学与技术设施委员会ISIS中子与μ子源)
AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models
AERMANI-VLM:基于视觉语言模型的空中 manipulation 的结构提示与推理
Sarthak Mishra, Rishabh Dev Yadav, Avirup Das, Saksham Gupta, Wei Pan, Spandan Roy
机构
*
Robotics Research Center, IIIT Hyderabad(IIIT海得拉巴机器人研究中心)
;
Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系)
;
Newcastle University(纽卡斯尔大学)