Selective Safety Steering via Value-Filtered Decoding
基于价值过滤解码的选择性安全引导
Bat-Sheva Einbinder, Hen Davidov, Yee Whye Teh, Yarin Gal, Yaniv Romano
机构
*
Department of Electrical and Computer Engineering, Technion IIT(技术学院电气与计算机工程系)
;
Department of Statistics, University of Oxford(牛津大学统计系)
;
OATML, Department of Computer Science, University of Oxford(牛津大学计算机科学系)
;
Department of Computer Science, Technion IIT(技术学院计算机科学系)
VehAnchor: Metadata-Free Metric Scale Recovery from Vehicle Cues in Aerial Imagery
VANGUARD:用于GPS受限环境下的无人机车辆锚定地面采样距离估计
Yifei Chen, Chenqian Le, Jiayi Cheng, Xupeng Chen
机构
*
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空航天信息研究所)
;
Tandon School of Engineering, New York University(纽约大学工学院)
;
School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院)
Comments@2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
Journal refProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 40785-40831 July 2-7, 2026
CommentsOral at the ICML 2026 Workshop on the Impact of Memorization on Trustworthy Foundation Models; Code available at https://github.com/Graph-COM/KVEraser
Safety from Honesty in a Disinterested AI Predictor
无兴趣AI预测器中的诚实安全性
Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gavenčiak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
机构
*
LawZero
;
Université de Montréal(蒙特利尔大学)
;
Mila
;
University of California Berkeley(加州大学伯克利分校)
;
McGill University(麦吉尔大学)
;
Arb Research
;
Center for Theoretical Study Charles University in Prague(布拉格查理大学理论研究中心)
;
University of Oxford(牛津大学)
;
Sapien Institute(Sapien研究所)
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,中国科学院自动化研究所)
;
Pengcheng Laboratory(鹏城实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Data Science and Artificial Intelligence Research Institute, China United Network Communications Group Co., Ltd.(中国联合网络通信集团有限公司数据科学与人工智能研究院)
DR-Arena: an Automated Evaluation Framework for Deep Research Agents
DR-Arena:深度研究智能体的自动化评估框架
Yiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan Zhang
机构
*
National University of Singapore(国立新加坡大学)
;
Nanyang Technological University(南洋理工大学)
;
Singapore Management University(新加坡管理学院)
;
Singapore University of Technology and Design(新加坡科技设计大学)
机构
*
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology (HKUST), Hong Kong SAR, China(香港科技大学计算机科学与工程系)
;
Tencent, Shenzhen, China(腾讯(中国深圳))
;
Shenzhen Institute of Advanced Technology (SIAT), Chinese Academy of Sciences, Shenzhen, China(深圳先进技术研究所(SIAT),中国科学院)