机构
*
Tsinghua University(清华大学)
;
Zhongguancun Academy(中关村学院)
;
Infinigence AI
;
Peking University(北京大学)
;
UC Berkeley(加州大学伯克利分校)
;
Harbin Institute of Technology(哈尔滨工程学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
CommentsAccepted to RSS 2026. This is the technical report of the RLinf Team, focusing on the algorithm side. For the system-level design, please refer to arXiv:2509.15965. The open-sourced code link: https://github.com/RLinf/RLinf
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
SteeringSafety:针对大语言模型在多安全视角下的表征引导基准测试
Vincent Siu, Nicholas Crispino, David Park, Nathan W. Henry, Zhun Wang, Yang Liu, Dawn Song, Chenguang Wang
机构
*
University of California, Santa Cruz(加州大学圣克鲁兹分校)
;
Washington University in St. Louis(华盛顿大学圣路易斯分校)
;
University of California, Berkeley(加州大学伯克利分校)