arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-07-20 至 2026-07-20 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2605.04539 2026-07-20 cs.CL cs.AI 版本更新 85%

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization

RLearner-LLM: 通过混合直接偏好优化平衡大语言模型中的逻辑性与流畅性

Qiming Bao, Juho Leinonen, Paul Denny, Michael J. Witbrock

机构 * University of Auckland(奥克兰大学) Aalto University(阿alto大学)

专题命中 偏好对齐 :RLHF(abstract,abstract_cn);DPO(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出RLearner-LLM,通过融合DeBERTa-v3 NLI信号与验证器LLM评分,解决直接偏好优化在知识密集型生成中的逻辑对齐问题,实现NLI指标显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01514 2026-07-20 cs.CL cs.AI cs.CR 版本更新 84%

Decoupled Alignment for Robust Plug-and-Play Adaptation

用于鲁棒即插即用适应的解耦对齐

Haozheng Luo, Jiahao Yu, Wenxin Zhang, Jialong Li, Chenghao Qiu, Yimin Wang, Eric Hanchen Jiang, Jerry Yao-Chieh Hu, Yan Chen, Binghui Wang, Xinyu Xing, Han Liu

机构 * Northwestern University(西北大学) New York University Abu Dhabi(纽约大学阿布扎克分校) Stanford University(斯坦福大学) Texas A&M University(德克萨斯农工大学) University of California, Los Angeles(加州大学洛杉矶分校) Illinois Institute of Technology(伊利诺伊理工学院)

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI

AI总结 研究提出无需训练的大语言模型对齐方法,利用知识蒸馏提取对齐信号,经模型融合实现即插即用的对齐校正,采用增量调试识别关键知识组件,在有害问题数据集上显著提升防御成功率,且不损性能。

Comments Revised to correct the Acknowledgments section. Previous versions inadvertently included acknowledgments of NSF and NIH awards that did not support this work. Those funding acknowledgments have been removed. The technical content, results, and conclusions are unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15095 2026-07-20 cs.CL cs.AI cs.MA 版本更新 79%

Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents

数字万神殿:使用大语言模型智能体模拟和审计联盟形成

Dylan Van Mulders, Matthias Bogaert, Dirk Van den Poel

机构 * Ghent University(根特大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI

AI总结 该研究针对政治联盟形成谈判,提出结合多种技术的多智能体框架,在弗拉芒选举中实施,通过引入相关拓扑、分数和检验使其可解释,模拟产生稳定结果,能可靠预测现实,提供了探索政党兼容性等的测试平台。

Comments 11 pages, 2 figures, 5 tables, To be published in the AIDEM Workshop proceedings of the ECML PKDD 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16117 2026-07-20 cs.CL 新提交 57%

Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content

语言编码的速率-效用前沿:在受控语言内容下比较词元、字节和像素

Ingo Ziegler, Martin Krebs, Desmond Elliott

机构 * Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

AI总结 研究在内容和下游能力受控时语言编码保留了什么,利用多语言平行句子通过共享瓶颈比较词元、字节和像素以追踪速率-效用前沿,评估三种效用,发现各编码在不同方面表现不同,选择编码需考虑多因素进行速率-效用权衡。

Comments Preprint. Code available at https://github.com/ziegler-ingo/rate-utility-frontiers

详情

展开后加载摘要…

URL PDF HTML 收藏