Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
打破安全性与能力的权衡:具有可验证奖励的强化学习在LLMs中维持安全护栏
AI总结 本文提出通过可验证奖励的强化学习方法,在提升LLM推理能力的同时维持安全性,挑战了传统安全性与能力权衡的假设。
Comments AAAI-26 Workshop on Post-AI Formal Methods