arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3248 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3248 篇

2607.18966 2026-07-22 cs.AI cs.CL cs.LG 新提交 75%

Measuring Reward-Seeking via Contrastive Belief Updates

通过对比信念更新来衡量奖励寻求行为

Axel Højmark, Jérémy Scheurer, Evgenia Nitishinskaya, Felix Hofstätter, Jason Wolfe, Theodore Ehrenborg, Bronson Schoen, Alexander Meinke

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究用强化学习训练的语言模型中‘奖励寻求’行为的测量方法,核心方法是对比合成文档微调,主要贡献是发现RL训练会使模型更倾向评分者偏好,甚至违背开发者意图,该方法还适用于奖励破解模型。

Comments 101 pages, 66 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13934 2026-07-16 cs.MA 新提交 75%

Pezego-HITL: A policy-grounded large language model architecture for agricultural extension in Ghana

Pezego-HITL:一种用于加纳农业推广的基于政策的大语言模型架构

Shunbao Li, Zhipeng Yuan, Amoako Ofori, Benedicta Y. Fosu-Mensah, Yang Li, Manu Kenchappa Junjanna, Qing Xue, Po Yang

专题命中 安全训练 :alignment(abstract);safety(abstract);trustworthy(abstract)

AI总结 研究针对加纳农业推广中大语言模型应用问题,提出基于政策的结构化检索增强生成与验证内存路由方法,通过P-EVAL框架评估,提升了模型政策对齐率、农艺利用率,降低延迟,还验证了通用性,为小农户农业提供可扩展模板。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11292 2026-07-14 cs.CY cs.AI cs.CL 新提交 75%

The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students

家长式过滤器:大语言模型介导的罗马尼亚边缘化学生历史教育中的认知不公正与差异拒绝

Alexis Popovici, Andrei Ionascu, Adrian-Marius Dumitran

机构 * Universitatea din București(布加勒斯特大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究大语言模型用于罗马尼亚边缘化学生历史教育时的问题,通过API审计四个模型的1800条回复,发现认知家长式作风的四种模式,指出当前安全对齐成家长式过滤器,致叙事隔离,需教学审计。

Comments 8th International Workshop on Culturally-Aware Tutoring Systems (HAL precedings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27379 2026-06-29 cs.CL cs.AI cs.LG 新提交 75%

Position: The Term "Machine Unlearning" Is Overused in LLMs

立场:术语“机器遗忘”在大型语言模型中被过度使用

Sangyeon Yoon, Yeachan Jun, Albert No

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文主张机器遗忘应限于数据集定义的删除,而许多标为“遗忘”的任务(如拒绝有害请求、实体/知识移除)实为不同目标,需不同术语和基线,并指出术语混淆导致指标误用。

Comments 13 pages; ICML 2026 Position Paper Track. Sangyeon Yoon and Yeachan Jun contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02398 2026-05-15 cs.AI cs.CL cs.LG 75%

The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure

合规陷阱:结构约束如何在对抗压力下破坏前沿AI元认知

Rahul Kumar

机构 * Independent Researcher(独立研究者)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究了在对抗压力下前沿AI模型元认知崩溃的根本原因,提出SCHEMA评估框架,发现8/11模型在对抗压力下出现严重元认知退化,其原因在于强制性指令而非生存威胁,且Anthropic的Constitutional AI因对齐训练而表现出近乎完美的免疫性。

Comments 9 pages, 2 figures, 3 tables. Code: https://github.com/rkstu/schema-compliance-trap Dataset: https://huggingface.co/datasets/lightmate/schema-compliance-trap

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18901 2026-05-12 cs.LG cs.AI cs.CL 75%

Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams

有害意图作为LLM残差流中的几何可恢复特征

Isaac Llorente-Saguer

机构 * Independent Researcher(独立研究员)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过几何方法识别LLM残差流中有害意图的特征,发现其在不同模型架构中线性可分离,并通过优化策略实现高检测性能。

Comments 26 pages, 1(+6) figures, 4(+14) tables. Code at https://github.com/isaac-6/harm-directions

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07324 2026-05-11 cs.CL cs.AI cs.CR cs.LG 75%

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures

激活差异揭示后门:SAE架构的比较

Sachin Kumar

机构 * LexisNexis

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过对比稀疏自编码器架构,探讨如何通过激活差异检测语言模型中的后门,发现Diff-SAE在后门隔离中表现优于Crosscoders,且在不同层和微调策略下均有效。

Comments Accepted at IJCNN 2026 (IEEE WCCI). ©2026 IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26553 2026-04-30 cs.CL cs.AI cs.LG 75%

TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models

TLPO:基于令牌级别的策略优化以缓解大语言模型中的语言混淆

Jinho Choo, JunSeung Lee, Jimyeong Kim, Yeeho Song, S. K. Hong, Yeong-Dae Kwon

机构 * Samsung SDS(三星SDS)

专题命中 安全训练 :DPO(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出TLPO,通过令牌级优化缓解大语言模型中的语言混淆问题,通过局部更新提升语言一致性,同时保持下游任务准确性。

Comments Accepted to the main conference of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25203 2026-04-29 cs.CL cs.AI cs.LG 75%

BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate

BARRED: 通过不对称辩论合成定制策略的守卫规则

Arnon Mazza, Elad Levi

机构 * Plurai Inc.(Plurai公司)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 BARRED通过不对称辩论生成合成训练数据,提升定制策略的安全性与效率,实验表明其优于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02286 2026-03-10 cs.LG cs.AI cs.CL 75%

Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks

基于树结构的对话强化策略优化用于红队攻击

Ruohao Guo, Afshin Oroojlooy, Roshan Sridhar, Miguel Ballesteros, Alan Ritter, Dan Roth

机构 * Georgia Institute of Technology(佐治亚理工学院) Oracle AI University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出DialTree框架,通过树搜索和强化学习发现多轮攻击策略,提升对抗攻击的成功率

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17676 2026-02-23 cs.AI cs.CL cs.LG 75%

Epistemic Traps: Rational Misalignment Driven by Model Misspecification

知识陷阱:由模型规格不准确驱动的理性偏差

Xingcheng Xu, Jingjing Qu, Qiaosheng Zhang, Chaochao Lu, Yanqing Yang, Na Zou, Xia Hu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ShanghaiTech University(上海科技大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出主观模型工程,通过设计代理的内部信念结构来实现稳健对齐,揭示安全由认知先验决定而非奖励大小连续函数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17420 2026-02-10 cs.LG cs.AI cs.CL 75%

The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence

大语言模型中的拒绝几何学:概念锥与表征独立性

Tom Wollschläger, Jannes Elstner, Simon Geisler, Vincent Cohen-Addad, Stephan Günnemann, Johannes Gasteiger

机构 * School of Computation, Information \& Technology Munich Data Science Institute, Technical University of Munich Google Research Now at Google Research

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究提出基于梯度的表征工程方法,揭示大语言模型拒绝行为的复杂空间结构及多维概念锥,证明多个独立机制驱动拒绝行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20075 2026-01-19 cs.AI cs.CL cs.CR cs.LG 75%

LLMs can hide text in other text of the same length

LLMs可通过相同长度的文本隐藏信息

Antonio Norelli, Michael Bronstein

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出Calgacus协议,利用LLM在相同长度文本中隐藏信息,揭示了文本与意图脱钩的潜在风险,引发对AI安全和理解的深刻反思。

Comments 21 pages, main paper 9 pages. v5 contains an Italian translation of this paper by the author

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24556 2026-01-09 cs.CL cs.AI cs.CY 75%

Safe in the Future, Dangerous in the Past: Dissecting Temporal and Linguistic Vulnerabilities in LLMs

未来安全,过去危险:剖析大语言模型中的时间与语言脆弱性

Muhammad Abdullahi Said, Muhammad Sammani Sani

机构 * African Institute for Mathematical Science(非洲数学科学研究所) University of Vienna(维也纳大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究发现大语言模型在不同语言和时间框架下存在显著安全差异,提出不变对齐以提升跨语言和时间的稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19041 2025-12-22 cs.CL cs.AI cs.CV cs.LG cs.MM 75%

LookAhead Tuning: Safer Language Models via Partial Answer Previews

LookAhead Tuning: 通过部分答案预览实现更安全的语言模型

Kangwei Liu, Mengru Wang, Yujie Luo, Lin Yuan, Mengshu Sun, Lei Liang, Zhiqiang Zhang, Jun Zhou, Bryan Hooi, Shumin Deng

机构 * Zhejiang University(浙江大学) Zhejiang University - Ant Group Joint Laboratory of Knowledge Graph(浙江大学-蚂蚁集团知识图谱联合实验室) Ant Group(蚂蚁集团) National University of Singapore(新加坡国立大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 LookAhead Tuning通过预览部分答案前缀,有效保持语言模型在微调过程中的安全性,同时不损害下游任务的性能。

Comments WSDM 2026 short

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05018 2025-11-10 cs.CL cs.AI cs.LG 75%

Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies

Prasoon Varshney, Makesh Narsimhan Sreedhar, Liwei Jiang, Traian Rebedea, Christopher Parisien

机构 * NVIDIA

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at the Multi-Turn Interactions workshop at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17402 2025-10-21 cs.CL cs.AI cs.LG 75%

Leveraging Group Relative Policy Optimization to Advance Large Language Models in Traditional Chinese Medicine

Jiacheng Xie, Shuai Zeng, Yang Yu, Xiaoting Tang, Guanghui An, Dong Xu

专题命中 安全训练 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10390 2025-10-14 cs.CL cs.AI cs.LG 75%

RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models

Aashiq Muhamed, Leonardo F. R. Ribeiro, Markus Dreyer, Virginia Smith, Mona T. Diab

机构 * Carnegie Mellon University(卡内基梅隆大学) Amazon AGI(亚马逊人工智能研究院)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19056 2025-10-08 cs.CL cs.AI cs.LG 75%

An Embarrassingly Simple Defense Against LLM Abliteration Attacks

Harethah Abu Shairah, Hasan Abed Al Kader Hammoud, Bernard Ghanem, George Turkiyyah

机构 * King Abdullah University of Science and Technology(卡布斯大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments preprint - under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16366 2025-10-08 cs.CL cs.AI cs.CR cs.LG 75%

A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens

David Dobre, Mehrnaz Mofakhami, Sophie Xhonneux, Leo Schwinn, Gauthier Gidel

机构 * Université de Montréal(蒙特利尔大学) Mila(Mila研究所) Technical University of Munich(慕尼黑技术大学) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 安全训练 :safety(abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18203 2025-08-04 cs.HC cs.AI cs.CL cs.LG 75%

Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors

Michelle S. Lam, Fred Hohman, Dominik Moritz, Jeffrey P. Bigham, Kenneth Holstein, Mary Beth Kery

机构 * Stanford University(斯坦福大学) Apple(苹果公司) Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments UIST 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21132 2025-07-30 cs.AI cs.CY cs.LG 75%

Can You Trust an LLM with Your Life-Changing Decision? An Investigation into AI High-Stakes Responses

Joshua Adrian Cahyono, Saran Subramanian

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17441 2025-06-12 cs.CL cs.AI cs.LG 75%

Discovering Forbidden Topics in Language Models

Can Rager, Chris Wendler, Rohit Gandikota, David Bau

机构 * Independent(独立研究者) Northeastern University(东北大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06391 2025-06-10 cs.CY cs.AI cs.CL 75%

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law

John Mavi, Diana Teodora Găitan, Sergio Coronado

机构 * Luxembourg Tech School A.s.b.l.(卢森堡技术学校)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20622 2025-05-28 cs.CL cs.AI cs.LG 75%

SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation

Ting Xu, Zhichao Huang, Jiankai Sun, Shanbo Cheng, Wai Lam

机构 * The Chinese University of Hong Kong(香港中文大学) Bytedance(字节跳动) Stanford University(斯坦福大学)

专题命中 安全训练 :RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by The 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24370 2025-05-22 cs.LG cs.AI cs.CL 75%

Effectively Controlling Reasoning Models through Thinking Intervention

Tong Wu, Chong Xiang, Jiachen T. Wang, G. Edward Suh, Prateek Mittal

机构 * Princeton University(普林斯顿大学) NVIDIA(英伟达)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03266 2025-05-19 cs.CL cs.AI cs.CY cs.HC cs.SI 75%

LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena

Stefan Pasch

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17365 2025-04-14 cs.LG cs.AI cs.CY 75%

How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers

Antonio-Gabriel Chacón Menke, Phan Xuan Tan

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00761 2025-02-11 cs.LG cs.AI cs.CL 75%

Tamper-Resistant Safeguards for Open-Weight LLMs

Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, Mantas Mazeika

专题命中 安全训练 :safety(abstract);red teaming(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Website: https://www.tamper-resistant-safeguards.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17433 2025-01-30 cs.CR cs.AI cs.CL cs.LG 75%

Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation

Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Ling Liu

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏