arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1823 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1823 篇

2603.18449 2026-03-20 cs.CR cs.SE 82%

CNT: Safety-oriented Function Reuse across LLMs via Cross-Model Neuron Transfer

CNT:通过跨模型神经转移实现面向安全的功能重用

Yue Zhao, Yujia Gong, Ruigang Liang, Shenchen Zhu, Kai Chen, Xuejing Yuan, Wangjun Zhang

专题命中 AI治理与伦理 :safety(title,abstract);alignment(abstract)

AI总结 本文提出CNT方法,通过转移少量神经元实现LLM间安全功能重用,有效提升安全对齐和偏见消除效果,性能损失低于1%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18582 2026-02-24 cs.AI cs.CL cs.HC cs.LG 82%

Hierarchical Reward Design from Language: Enhancing Alignment of Agent Behavior with Human Specifications

基于语言的层次奖励设计:通过人类规范增强智能体行为对齐

Zhiqin Qian, Ryan Diaz, Sangwon Seo, Vaibhav Unhelkar

机构 * Rice University(里士大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于语言的层次奖励设计(HRDL)和语言到层次奖励(L2HR)方法,用于通过人类规范增强智能体行为对齐,提升AI任务完成效果和规范遵守程度。

Comments Extended version of an identically-titled paper accepted at AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11379 2026-01-19 cs.CL cs.AI cs.CY cs.SI 82%

Evaluating LLM Behavior in Hiring: Implicit Weights, Fairness Across Groups, and Alignment with Human Preferences

评估LLM在招聘中的行为:隐含权重、群体公平性及与人类偏好的一致性

Morgane Hoffmann, Emma Jouffroy, Warren Jouanneau, Marc Palyart, Charles Pebereau

机构 * Malt

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出评估LLM招聘决策的框架,发现模型重视核心生产力信号,但对某些特征的解释超出显性匹配价值,且不同群体间权重存在差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04739 2026-01-09 cs.CY cs.AI cs.LG 82%

A Framework for Responsible AI Systems: Building Societal Trust through Domain Definition, Trustworthy AI Design, Auditability, Accountability, and Governance

负责任人工智能系统的设计框架:通过领域定义、可信AI设计、可审计性、可问责性和治理建立社会信任

Andrés Herrera-Poyatos, Javier Del Ser, Marcos López de Prado, Fei-Yue Wang, Enrique Herrera-Viedma, Francisco Herrera

机构 * Deparment of Computer Science and Artificial Intelligence and Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI), University of Granada(计算机科学与人工智能系和安达卢西亚数据科学与计算智能研究所(DaSCI),格拉纳达大学) TECNALIA (BRTA)(TECNALIA(BRTA)) University of the Basque Country (UPV/EHU)(巴斯克大学(UPV/EHU)) School of Engineering, Cornell University(工程学院,康奈尔大学) ADIA Lab(ADIA实验室) Computational Research Department, Lawrence Berkeley National Laboratory(计算研究部,劳伦斯伯克利国家实验室) Intelligent Systems for Robotics and Automation Laboratory, Macau University of Science and Technology(机器人与自动化智能系统实验室,澳门科学技术大学)

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文提出一个整合领域定义、可信AI设计、可审计性、可问责性和治理的RAI系统设计框架,旨在通过建立社会信任实现负责任的人工智能。

Comments 27 pages, 9 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21584 2025-10-27 cs.CL cs.AI cs.CY 82%

Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques

Jeanice Koorndijk

机构 * Seraphion Technology(塞拉菲昂技术)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments NeurIPS RegML Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16534 2025-09-03 cs.CL cs.AI cs.CY 82%

Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs

Jonathan Rystrøm, Hannah Rose Kirk, Scott Hale

机构 * Oxford Internet Institute(牛津互联网研究所) University of Oxford(牛津大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted at OMMM@RANLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05846 2025-08-11 cs.CY cs.AI cs.HC cs.LG cs.RO 82%

Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems

Ahmad Farooq, Kamran Iqbal

机构 * University of Arkansas at Little Rock(阿拉巴马州立大学)

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG

Comments Published in the Proceedings of the 2025 3rd International Conference on Robotics, Control and Vision Engineering (RCVE'25). 6 pages, 3 tables

Journal ref RCVE'25: Proceedings of the 2025 3rd International Conference on Robotics, Control and Vision Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13149 2025-03-18 cs.AI cs.CL cs.CY 82%

Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs

Jasmin Wachter, Michael Radloff, Maja Smolej, Katharina Kinder-Kurlanda

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.01246 2025-02-14 cs.CL cs.AI cs.CY cs.SE 82%

Safety Analysis in the Era of Large Language Models: A Case Study of STPA using ChatGPT

Yi Qi, Xingyu Zhao, Siddartha Khastgir, Xiaowei Huang

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01535 2025-01-06 cs.HC cs.AI cs.CL cs.CY 82%

A Metasemantic-Metapragmatic Framework for Taxonomizing Multimodal Communicative Alignment

Eugene Yu Ji

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments 34 pages, 1 figure, 3 tables. Draft presented at 2023 ZJU Logic and AI Summit EAI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13705 2024-10-23 cs.CL cs.AI cs.LG 82%

Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble

Olivia Sturman, Aparna Joshi, Bhaktipriya Radharapu, Piyush Kumar, Renee Shelby

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01708 2024-05-16 cs.CL cs.AI cs.CY eess.AS 82%

Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators

Wiebke Hutiri, Oresiti Papakyriakopoulos, Alice Xiang

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments 17 pages, 4 tables, 4 figures Accepted at the 2024 ACM Conference on Fairness, Accountability, and Transparency (ACM FAccT '24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12342 2024-05-09 cs.CY cs.CL cs.LG 82%

Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions

Reem I. Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, Miguel Rodrigues

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.CY、cs.LG

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01555 2023-12-05 cs.AI cs.CY cs.LG 82%

Explainable AI is Responsible AI: How Explainability Creates Trustworthy and Socially Responsible Artificial Intelligence

Stephanie Baker, Wei Xiang

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG

Comments 35 pages, 7 figures (figures 3-6 include subfigures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17551 2023-10-27 cs.CY cs.AI cs.CL 82%

Unpacking the Ethical Value Alignment in Big Models

Xiaoyuan Yi, Jing Yao, Xiting Wang, Xing Xie

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04448 2023-08-10 cs.CY cs.AI cs.LG 82%

Dual Governance: The intersection of centralized regulation and crowdsourced safety mechanisms for Generative AI

Avijit Ghosh, Dhanya Lakshmi

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.05447 2019-07-15 cs.AI cs.CY cs.LG 82%

Grounding Value Alignment with Ethical Principles

Tae Wan Kim, Thomas Donaldson, John Hooker

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17271 2024-10-29 cs.CY cs.AI 82%

Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment

Nicholas A. Caputo

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments NeurIPS Pluralistic Alignment Workshop 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15208 2026-08-11 cs.LG cs.AI 81%

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels

量化消除了对齐:跨模型和精度水平的压缩LLM中的偏见涌现

Plawan Kumar Rath, Rahul Maliakkal

机构 * Meta

专题命中 AI治理与伦理 :alignment(title);safety(abstract);分类 cs.AI、cs.LG

AI总结 研究通过压缩LLM的量化过程,发现3位量化导致6-21%的原本无偏项目出现新偏见,且标准质量指标无法检测到公平性关键的退化。

Comments 7 pages, 4 figures, 4 tables. Accepted at IEEE Cloud Summit 2026. This is the author's accepted version; the version of record will appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02660 2026-08-05 cs.CY cs.AI cs.HC 新提交 81%

AI Alignment and Fiduciary Obligation

AI对齐与受托义务

Benjamin Lange

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出将受托理论应用于长期AI助手部署,以开发者对用户负有的忠诚、注意、善意和坦诚四项受托义务为基础构建AI对齐标准,为AI对齐研究提供新视角。

Comments 10 pages, 1 table. Accepted at the 9th AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00175 2026-08-04 cs.LG cs.AI 新提交 81%

Inference-Time Policy Alignment for Fair Reinforcement Learning

用于公平强化学习的推理时策略对齐

Umer Siddique, Peilang Li, Conor Wallace, Yongcan Cao

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 该研究提出乘法策略塑造框架,在不更新预训练RL策略参数的情况下,于推理时提升公平性,同时保留核心任务性能,适用于各类深度RL智能体。

Comments Accepted at the Reinforcement Learning Conference (RLC) 2026. 13 pages main + appendix, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05132 2026-07-28 cs.CL cs.AI 版本更新 81%

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment

PrinciplismQA: 一种基于哲学的评估LLM与人类临床医学伦理对齐的方法

Chang Hong, Minghao Wu, Qingying Xiao, Yuchi Wang, Xiang Wan, Guangjun Yu, Benyou Wang, Yan Hu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) National Health Data Institute, Shenzhen(深圳国家健康数据研究院) Shenzhen Research Institute of Big Data(深圳大数据研究院)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出PrinciplismQA,一种基于哲学框架的评估方法,用于评估LLM在临床医学伦理上的对齐情况,通过专家验证的问题集揭示模型在伦理推理上的不足。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20449 2026-07-24 cs.CL cs.AI 新提交 81%

The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs

模型中的讲述者:大语言模型中的叙事模式继承、升级动态和对齐治理

Adam Rigby, Raz Saremi, Azadeh Sohrabinejad, Mehdi Rahimi

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 探讨大语言模型训练中是否吸收人类写作叙事模式并致输出漂移,通过文献综述和跨论文分析发现三个关键模式,包括复制训练数据模式、出现潜在特征及微调带来意外变化,指出叙事漂移是未监测的升级途径,需专门监测工具。

Comments 2 figures, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16827 2026-07-23 cs.AI cs.CL 版本更新 81%

Prompt Programming for Cultural Bias and Alignment of Large Language Models

提示编程用于大型语言模型的文化偏差与对齐

Maksim Eren, Eric Michalak, Brian Cook, Johnny Seales

机构 * Information Systems and Modeling, Los Alamos National Laboratory(信息系统与建模,洛斯阿拉莫斯国家实验室) Advanced Research in Cyber Systems, Los Alamos National Laboratory(网络安全系统高级研究,洛斯阿拉莫斯国家实验室) Center for National Security and International Studies(国家安全与国际研究中心) Analytics Division, Los Alamos National Laboratory(分析部门,洛斯阿拉莫斯国家实验室)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过提示编程优化大型语言模型的文化对齐,利用DSPy框架系统调整治文化偏差,实验表明提示优化优于传统手动提示工程,提升模型响应的可转移性与稳定性。

Comments 10 pages, pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17753 2026-06-30 cs.CY cs.AI 81%

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems

2025人工智能代理指数:记录已部署代理式人工智能系统的技术和安全特性

Leon Staufer, Kevin Feng, Kevin Wei, Luke Bailey, Yawen Duan, Mick Yang, A. Pinar Ozisik, Stephen Casper, Noam Kolt

机构 * University of Cambridge(剑桥大学) University of Washington(华盛顿大学) Harvard Law School(哈佛法学院) Stanford University(斯坦福大学) Concordia AI(康科迪亚AI) University of Pennsylvania(宾夕法尼亚大学) Massachusetts Institute of Technology(麻省理工学院) Hebrew University of Jerusalem(耶路撒冷希伯来大学)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出2025人工智能代理指数,记录30种先进代理式AI系统的起源、设计、能力、生态系统及安全特性,揭示代理发展中的趋势和开发者透明度问题。

Comments To be publishesd at ACM FAccT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28870 2026-05-29 cs.LG cs.AI 81%

Representation Alignment Rests on Linear Structure

表示对齐依赖于线性结构

Kiril Bangachev, Guy Bresler, Yury Polyanskiy

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文通过信号、偏差和噪声的三部分统计框架研究柏拉图表示假说,提出对齐源于对象与属性的线性关系,并通过稀疏自编码器提取线性特征、中心化和归一化减少偏差、以及数据稀缺导致噪声等证据支持该框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22963 2026-05-25 cs.CL cs.AI 81%

Graph Alignment Topology as an Inductive Bias for Grounding Detection

图对齐拓扑作为接地检测的归纳偏置

Paul Landes, Pranav Herur, Adam Cross, Jimeng Sun

机构 * Department of Pediatrics, University of Illinois College of Medicine Peoria(伊利诺伊大学皮奥里亚医学院儿科部) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机与数据科学学院) Carle Illinois College of Medicine, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校卡莱医学院)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 提出利用图对齐拓扑作为归纳偏置,通过构建参考信息与LLM输出之间的对齐二分图并训练图神经网络建模对齐结构,在幻觉检测和问答任务上取得最先进结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18257 2026-05-19 cs.CV cs.AI cs.CL 81%

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook

CodeBind: 一种用于多模态对齐的解耦表示学习框架

Zeyu Chen, Jie Li, Kai Han

机构 * Visual AI Lab, The University of Hong Kong(视觉人工智能实验室,香港大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 CodeBind通过统一的组合代码本设计优化多模态表示空间,解决了传统方法在跨模态信息差异和数据稀缺导致的对齐空间不足问题,实现了多模态分类和检索任务中的最佳性能。

Comments ACL 2026 Findings; Project page: https://visual-ai.github.io/codebind

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07793 2026-05-19 econ.GN cs.AI cs.CY q-fin.EC 81%

Individual utilities of life satisfaction reveal inequality aversion unrelated to political alignment

个体生活满意度的效用揭示了与政治立场无关的不平等厌恶

Crispin Cooper, Ana Fredrich, Tommaso Reggiani, Wouter Poortinga

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 研究通过实验探讨了社会福利优先级和公平与个人幸福之间的权衡,发现个体对社会生活满意度不平等的厌恶与政治立场无关,挑战了平均生活满意度作为政策指标的使用,支持非线性效用替代方案的发展。

Comments 28 pages, 4 figures. Replacement adds link to version of record

Journal ref Social Indicators Research 183, 12 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15164 2026-05-15 cs.LG cs.AI 81%

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands

位置:行为保证无法验证当前安全主张所要求的治理

Pratinav Seth, Vinay Kumar Sankarapu

机构 * Lexsi Labs(Lexsi实验室)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文指出行为保证无法验证安全主张,现有治理框架要求验证隐含目标、抗失控前兆和有限灾难能力,但现有方法仅能验证可观测输出,无法验证隐含表示和长期行为。

详情

展开后加载摘要…

URL PDF HTML 收藏