Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
少些推理,多些验证:确定性门控在使用工具的语言模型智能体中恢复了一种无声的违反策略失败模式
机构 * Indian Institute of Technology Kharagpur(印度理工学院卡拉格布尔分校) ; Massachusetts Institute of Technology(麻省理工学院)
AI总结 研究使用工具的语言模型智能体违反策略问题,提出用确定性预执行门控干预,在τ²基准航空公司领域评估,该方法能提高成功率,防止无声违反策略写入,虽不保证任务成功,但有可靠性成果。