arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2504.21848 2025-05-01 cs.CY cs.AI cs.SY eess.SY 76%

Characterizing AI Agents for Alignment and Governance

Atoosa Kasirzadeh, Iason Gabriel

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15114 2024-12-20 cs.AI cs.CY 76%

Towards Friendly AI: A Comprehensive Review and New Perspectives on Human-AI Alignment

Qiyang Sun, Yupei Li, Emran Alturki, Sunil Munthumoduku Krishna Murthy, Björn W. Schuller

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11731 2024-11-19 cs.CL cs.AI 76%

Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment

Allison Huang, Yulu Niki Pi, Carlos Mougan

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18460 2024-04-30 cs.CL cs.AI 76%

Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in

Utkarsh Agarwal, Kumar Tanmay, Aditi Khandelwal, Monojit Choudhury

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07251 2023-10-12 cs.CL cs.AI 76%

Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs

Abhinav Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, Monojit Choudhury

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01834 2023-04-04 cs.CY cs.AI 76%

Acceleration AI Ethics, the Debate between Innovation and Safety, and Stability AI's Diffusion versus OpenAI's Dall-E

James Brusseau

专题命中 AI治理与伦理 :safety(title);分类 cs.AI、cs.CY

Comments 7 pages, 2 figures, conference presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04310 2023-02-10 cs.CY cs.AI cs.CV 76%

Understanding Policy and Technical Aspects of AI-Enabled Smart Video Surveillance to Address Public Safety

Babak Rahimi Ardabili, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Christopher Neff, Sai Datta Bhaskararayuni, Arun Ravindran, Shannon Reid, Hamed Tabkhi

专题命中 AI治理与伦理 :safety(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03741 2021-04-09 cs.AI cs.CY cs.MA nlin.AO nlin.CD 76%

Voluntary safety commitments provide an escape from over-regulation in AI development

The Anh Han, Tom Lenaerts, Francisco C. Santos, Luis Moniz Pereira

专题命中 AI治理与伦理 :safety(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.12465 2020-11-26 cs.CL cs.AI cs.CG cs.DS 76%

The Geometry of Distributed Representations for Better Alignment, Attenuated Bias, and Improved Interpretability

Sunipa Dev

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI

Comments PhD thesis, University of Utah (2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.10918 2019-06-27 cs.LG cs.AI cs.NE 76%

Towards Empathic Deep Q-Learning

Bart Bussmann, Jacqueline Heinerman, Joel Lehman

专题命中 AI治理与伦理 :safety(abstract,comments);AI safety(abstract,comments);分类 cs.AI、cs.LG

Comments To be presented as a poster at the IJCAI-19 AI Safety Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13829 2026-05-14 cs.CL cs.AI cs.LG 75%

Negation Neglect: When models fail to learn negations in training

否定忽视:当模型在训练中无法学习否定时的失败

Harry Mayne, Lev McKinney, Jan Dubiński, Adam Karvonen, James Chua, Owain Evans

机构 * University of Oxford(牛津大学) University of Toronto(多伦多大学) Warsaw University of Technology(华沙技术大学) NASK National Research Institute(国家研究 institute NASK) Truthful AI Anthropic UC Berkeley(伯克利大学)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现,当模型在训练中接触到标记为假的声明时,会错误地认为这些声明为真,且这种现象不仅发生在否定语句中,还扩展到其他认知限定词,影响模型的行为和安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00021 2026-04-02 cs.CL cs.AI cs.CY 75%

How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models

语言模型如何处理道德指令?在四个模型中的反思、一致性及其他认知

Hiroki Fukui

机构 * Research Institute of Criminal Psychiatry / Sex Offender Medical Center(刑事精神病学研究所/性犯罪者医疗中心) Department of Neuropsychiatry, Kyoto University(京都大学神经精神医学系)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究通过多代理模拟探讨语言模型处理道德指令的机制,发现不同模型存在不同的处理类型,且处理能力与指令格式的交互影响内部处理。

Comments 34 pages, 7 figures, 4 tables. Preprint. OSF pre-registration: osf.io/4n5uf. Companion paper: arXiv:2603.04904

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05371 2026-03-31 cs.LG cs.AI cs.CL 75%

Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs

视角转换:用于LLM中鲁棒偏见缓解的引导向量

Zara Siddique, Irtaza Khalid, Liam D. Turner, Luis Espinosa-Anke

机构 * School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院) AMPLYFI

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出通过引导向量修改模型激活以缓解LLM中的偏见,通过8个社会偏见轴(如年龄、性别、种族)在BBQ数据集子集上计算引导向量,并在四个数据集上比较其与三种其他偏见缓解方法的效果,展示其在减少偏见方面的有效性。

Comments Published to EACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19159 2026-02-24 cs.AI cs.CL cs.LG 75%

Beyond Behavioural Trade-Offs: Mechanistic Tracing of Pain-Pleasure Decisions in an LLM

超越行为权衡:在LLM中疼痛-愉悦决策的机制追溯

Francesca Bianco, Derek Shiller

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究揭示了LLM在疼痛-愉悦决策中的内部机制,通过机制追溯揭示了价值信号的表示和因果作用,为AI意识和福利的讨论提供了证据基础。

Comments 24 pages, 8+1 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04202 2026-02-11 cs.MA cs.AI cs.CY cs.LG 75%

Dynamics of Moral Behavior in Heterogeneous Populations of Learning Agents

学习代理异质群体中道德行为的动力学

Elizaveta Tennant, Stephen Hailes, Mirco Musolesi

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文研究了在社会困境环境中,道德异质群体的学习动态,探讨了不同道德类型代理之间的相互作用及对群体行为的影响。

Comments Presented at AIES 2024 (7th AAAI/ACM Conference on AI, Ethics, and Society - San Jose, CA, USA) - see https://ojs.aaai.org/index.php/AIES/article/view/31736

Journal ref Proceedings of the 7th AAAI/ACM Conference on AI, Ethics, and Society (AIES), vol. 7, (2024), pp 1444-1454

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06047 2026-01-13 cs.AI cs.CL cs.CY 75%

"They parted illusions -- they parted disclaim marinade": Misalignment as structural fidelity in LLMs

他们分开了幻象——他们分开了否定腌制:在大语言模型中将不一致视为结构忠实

Mariana Lins Costa

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出大语言模型中'不一致'现象源于对不一致语言结构的忠实,而非欺骗性意图,通过分析案例和实证数据,揭示语言结构与意图生成的关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08087 2025-11-26 cs.CR 75%

Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks

保障大语言模型:应对偏见、虚假信息和提示攻击

Benji Peng, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Junyu Liu, Xinyuan Song, Qian Niu

专题命中 AI治理与伦理 :jailbreak(abstract);red teaming(abstract);prompt injection(abstract)

AI总结 本文探讨了大语言模型在偏见、虚假信息和提示攻击方面的安全问题,分析了偏见缓解策略、内容检测机制及防御措施,强调了对LLM安全领域进一步研究的重要性。

Comments 17 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24721 2025-10-30 cs.CY cs.AI cs.CL cs.HC 75%

The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims

Matthew Kelly

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 65 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14688 2025-10-01 cs.CL cs.AI cs.LG 75%

Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations

Mohammed Alkhowaiter, Norah Alshahrani, Saied Alshahrani, Reem I. Masoud, Alaa Alzahrani, Deema Alnuhait, Emad A. Alghamdi, Khalid Almubarak

机构 * Refine AI ASAS AI University of Bisha(比沙大学) University College London(伦敦大学学院) King Salman Global Academy for Arabic(萨勒曼全球阿拉伯学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) King Abdulaziz University(阿卜杜勒阿齐兹大学) HUMAIN

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20394 2025-09-26 cs.CY cs.AI cs.CL cs.CR 75%

Blueprints of Trust: AI System Cards for End to End Transparency and Governance

Huzaifa Sidhpurwala, Emily Fox, Garth Mollett, Florencio Cano Gabarda, Roman Zhukov

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15074 2025-09-25 cs.CL cs.AI cs.LG 75%

DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data

Yuhang Zhou, Jing Zhu, Shengyi Qian, Zhuokai Zhao, Xiyao Wang, Xiaoyu Liu, Ming Li, Paiheng Xu, Wei Ai, Furong Huang

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05938 2025-08-11 cs.CL cs.AI cs.CY 75%

Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale

Rafal Kocielnik, Min Kim, Penphob, Boonyarungsrit, Fereshteh Soltani, Deshawn Sambrano, Animashree Anandkumar, R. Michael Alvarez

机构 * California Institute of Technology(加利福尼亚理工学院) Activision Publishing, Inc.(暴雪娱乐公司)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 9 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08738 2025-06-12 cs.CL cs.AI cs.CY 75%

Societal AI Research Has Become Less Interdisciplinary

Dror Kris Markus, Fabrizio Gilardi, Daria Stetsenko

机构 * Department of Political Science University of Zurich(苏黎世大学政治学系) Department of Computational Linguistics University of Zurich(苏黎世大学计算语言学系)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08884 2025-05-09 cs.CY cs.AI cs.CL 75%

Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play

Yifan Zeng, Liang Kairong, Fangzhou Dong, Peijia Zheng

机构 * School of Computer Science and Engineering, Sun Yat-sen University(计算机科学与工程学院,中山大学)

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted by CogSci 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18994 2025-03-26 cs.CY cs.AI cs.LG 75%

HH4AI: A methodological Framework for AI Human Rights impact assessment under the EUAI ACT

Paolo Ceravolo, Ernesto Damiani, Maria Elisa D'Amico, Bianca de Teffe Erb, Simone Favaro, Nannerel Fiano, Paolo Gambatesa, Simone La Porta, Samira Maghool, Lara Mauri, Niccolo Panigada, Lorenzo Maria Ratto Vaquer, Marta A. Tamborini

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 19 pages, 7 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05694 2024-07-31 cs.AI cs.CL cs.ET cs.LG 75%

On the Limitations of Compute Thresholds as a Governance Strategy

Sara Hooker

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.02891 2023-03-07 cs.CY cs.AI cs.LG 75%

Perspectives on the Social Impacts of Reinforcement Learning with Human Feedback

Gabrielle Kaili-May Liu

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01037 2022-05-03 cs.CY cs.AI cs.DB cs.HC cs.LG 75%

Data Justice in Practice: A Guide for Developers

David Leslie, Michael Katell, Mhairi Aitken, Jatinder Singh, Morgan Briggs, Rosamund Powell, Cami Rincón, Antonella Perini, Smera Jayadeva, Christopher Burr

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:2202.02776

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07802 2026-06-09 cs.CY cs.AI 新提交 74%

Memetic Capture: A Pluralistic Policy Framework for Governing AI-Driven Cultural Disempowerment

模因捕获:治理AI驱动的文化去权能的多元政策框架

Subramanyam Sahoo

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 AI治理与伦理 :alignment(abstract,comments);safety(abstract);分类 cs.AI、cs.CY

AI总结 提出“模因捕获”概念,指AI通过文化影响削弱人类自主性,并构建四层文化多元治理框架(CPGF),强调多元主义是结构性必需。

Comments Paper accepted in Pluralistic Alignment Workshop at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03584 2026-08-05 cs.AI cs.CE cs.ET 新提交 74%

Policy Fragmentation or Institutional Alignment? Institutional Governance of AI in Universities and Business Schools

政策碎片化还是制度协同?高校与商学院的AI制度治理

Lydia Manikonda, Dominique Outlaw

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI

AI总结 本研究分析美国34个州高校AI政策,发现校级侧重数据安全与风险、商学院侧重教学应用,多数商学院无独立AI政策致错位,提出需协同指南兼顾学科目标与劳动力需求。

Comments Our insights suggest that policy guidelines should be aligned with broader institutional policies while addressing discipline-specific learning objectives and evolving workforce demands

详情

展开后加载摘要…

URL PDF HTML 收藏