arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-08-19 至 2026-08-19 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 4 篇

2602.17127 2026-08-19 cs.CL 版本更新 74%

The Emergence of Lab-Driven Alignment Signatures: A Psychometric Framework for Auditing Latent Bias and Compounding Risk in Generative AI

实验室驱动的对齐签名的出现:一种心理测量框架,用于审计生成AI中的潜在偏见和叠加风险

Dusan Bosnjakovic

机构 * AI Researcher(人工智能研究员)

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL

AI总结 本文提出了一种心理测量框架,用于审计生成AI中的潜在偏见和叠加风险,通过分析九个领先模型的实验室信号,揭示了持续行为聚类的成因。

Comments v2: expanded from 9 to 18 behavioral dimensions and from 4 to 6 developer organizations; revised statistical methodology (rank-based inference with effect-size criterion, replacing variance-decomposition approach); model-level results now reported; references corrected throughout

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16893 2026-08-19 cs.CY cs.AI 新提交 62%

A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications

在安全调查中使用和评估大语言模型(LLM)作为代理专家的框架:可靠性、偏差及启示

Despoina Giarimpampa, Roland Meier, Tegawendé F. Bissyandé, Vincent Lenders, Jacques Klein

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究提出了评估LLM作为安全调查代理专家的框架,发现LLM虽内部一致但与专家响应存在系统性偏差,可用于试点和假设生成但不能替代专家征询。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16891 2026-08-19 cs.AI cs.CE cs.CR cs.CY 新提交 62%

Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution

智能体AI的运行时治理:基于可信溯源与故障闭锁执行的行动边界控制

Adam Mazzocchetti

机构 * SPQR Technologies Inc.(SPQR科技公司)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 该研究提出Aegis运行时治理系统,通过可信决策层调解智能体AI的工具行动提案,在沙堡语料库评估中成功阻止风险提案转化为治理副作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17516 2026-08-19 cs.CL 新提交 57%

Effects of Answer Format Variation on Gender Bias in Large Language Models

答案格式变化对大型语言模型中性别偏见的影响

Ksenia Merzlyakova, Sebastian Padó, Franziska Weeber

机构 * Institute for Natural Language Processing, University of Stuttgart(斯图加特大学自然语言处理研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 该研究探讨答案格式变化对LLMs性别偏见测量的影响,通过评估三种指令微调模型在不同格式下的表现,发现格式会显著改变测量结果,强调需将答案格式纳入LLM评估。

Comments 6th Workshop on Computational Linguistics for the Political and Social Sciences (CPSS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏