机构
*
School of Cyber Science and Technology, Beihang University, Beijing, China(北京航空航天大学网络安全学院)
;
Institute of Artificial Intelligence, Beihang University, Beijing, China(北京航空航天大学人工智能研究院)
;
University of Chinese Academy of Sciences, Beijing, China(中国科学院大学)
;
AI Security Lab, Beijing, China(360人工智能安全实验室)
Selective Safety Steering via Value-Filtered Decoding
基于价值过滤解码的选择性安全引导
Bat-Sheva Einbinder, Hen Davidov, Yee Whye Teh, Yarin Gal, Yaniv Romano
机构
*
Department of Electrical and Computer Engineering, Technion IIT(技术学院电气与计算机工程系)
;
Department of Statistics, University of Oxford(牛津大学统计系)
;
OATML, Department of Computer Science, University of Oxford(牛津大学计算机科学系)
;
Department of Computer Science, Technion IIT(技术学院计算机科学系)
CLAR: Learning 3D Representations for Robotic Manipulation by Fusing Masked Reconstruction with Multi-Level Contrastive Alignment
CLAR: 通过融合掩码重建与多层级对比对齐学习用于机器人操作的3D表示
Wenbo Cui, Chengyang Zhao, Yuhui Chen, Haoran Li, Zhizheng Zhang, Dongbin Zhao, He Wang
机构
*
SKL-MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所SKL-MAIS)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Galbot
;
CFCS, School of Computer Science, Peking University(北京大学计算机科学与技术学院CFCS)
CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs
CausalT5k: 诊断可信因果推理中的拒绝与失败模式——跨越因果阶梯
Longling Geng, Andy Ouyang, Theodore Wu, Daphne Barretto, Matthew John Hayes, Rachael Cooper, Yuqiao Zeng, Sameer Vijay, Gia Ancone, Ankit Rai, Matthew Wolfman, Patrick Flanagan, Edward Y. Chang
机构
*
Zhejiang University(浙江大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
Hong Kong University of Science and Technology (GZ)(香港科技大学(广州))
;
Nanjing University(南京大学)
BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment
评估两用生物学环境中的校准拒绝和安全有用性
Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, Kenny Workman
Ming Wen, Yuxuan Liu, Kun Yang, Yunhao Feng, Zhuoer Xu, Yuhao Sun, Shiwen Cui, Xiang Zheng, Yi Liu, Xingjun Ma, Yu-Gang Jiang
机构
*
Institute of Trustworthy Embodied AI(可信具身人工智能研究院)
;
Fudan University(复旦大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
Ant Group(蚂蚁集团)
;
Zhejiang University(浙江大学)
;
City University of Hong Kong(香港城市大学)
TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech
TANDEM: 面向多模态仇恨言论的时间感知神经检测
Girish A. Koushik, Helen Treharne, Diptesh Kanojia
机构
*
Nature-Inspired Computing & Engineering, University of Surrey(Surrey大学自然启发计算与工程系)
;
Surrey Centre for Cyber Security, University of Surrey(Surrey大学网络安全中心)
CommentsThis submission has been withdrawn because it has been superseded by a substantially revised and expanded version, available as arXiv:2607.27373. Please refer to and cite the newer version