arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-08-14 至 2026-08-14 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 4 篇

2608.13317 2026-08-14 cs.AI 新提交 74%

StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems

StateBridge:面向大语言模型多智能体系统潜在通信的无训练隐藏状态对齐

Yanwen Peng, Delvin Ce Zhang, Xi Wang, Nikolaos Aletras

机构 * School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院)

专题命中 幻觉与事实性 :alignment(title);分类 cs.AI

AI总结 StateBridge 是一种无训练的潜在通信方法,通过闭式正交变换对齐大语言模型多智能体的隐藏状态,在 26 个模型-任务对中 22 个取得最优或并列最优性能,优于基线。

Comments 18 pages, 3 figures, 4 tables, accepted by COLM2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14390 2026-08-14 cs.CL cs.AI 版本更新 62%

REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Models

REHEARSE:大语言模型中用于语言置信度校准的经验式排练

Ke Fang, Tianyi Zhao, Qianwen Wang, Lu Cheng

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Southern California(南加州大学) University of Illinois Chicago(伊利诺伊大学香槟分校)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

AI总结 针对大语言模型置信度与实际正确性不匹配的问题,提出无训练的Rehearse方法,通过置信度校准博弈的反馈生成校准信号,在多模型多基准实验中显著降低了预期校准误差并提升准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23982 2026-08-14 cs.MA cs.AI 版本更新 57%

Moral Hazard in Multi-Agent Language Models

多智能体语言模型中的道德风险

Dane Malenfant

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

AI总结 研究多智能体语言模型中的道德风险,引入对话道德风险游戏,评估七个模型并分解行为,用多种优化机制更新,发现效果异质,强调应报告机制级行为而非仅团队成功。

Comments This revision substantially expands the empirical evaluation to eleven open-weight and three frontier models, adding matched query-cost, team-reward, group-size, and private-share incentive analyses. It also extends the weight-level and GEPA results, frozen-prompt information-structure interventions, statistical uncertainty analyses, and qualitative prompt/reasoning-trace studies

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13658 2026-08-14 cs.LG 版本更新 57%

Post-Hoc Uncertainty-Aware Explanations for Deployed Power Quality Disturbance Classifiers via Laplace Approximation

基于不确定性的贝叶斯解释框架用于电力质量问题分类

Yinsong Chen, Samson S. Yu, Kashem M. Muttaqi

机构 * School of Engineering, Deakin University(德肯大学工程学院) ARC Training Centre in Energy Technologies for Future Grids, School of Engineering, University of Wollongong(未来电网能源技术培训中心,沃尔灵宗大学工程学院)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

AI总结 本文提出一种贝叶斯解释框架,通过生成相关性归因分布来建模解释不确定性,提升电力质量问题分类器的透明度和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏