arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1823 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1823 篇

2512.03047 2025-12-04 cs.CL cs.AI cs.LG 87%

Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models

基于熵的大型语言模型价值漂移与对齐工作的测量

Samih Fadli

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于熵的框架,用于测量大型语言模型中的价值漂移和对齐工作,通过五种行为分类法和熵动态分析,实现对模型伦理熵的实时监控与警报。

Comments 6 pages. Companion paper to "The Second Law of Intelligence: Controlling Ethical Entropy in Autonomous Systems". Code and tools: https://github.com/AerisSpace/EthicalEntropyKit

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12088 2025-07-29 cs.CR cs.CY 86%

Risks & Benefits of LLMs & GenAI for Platform Integrity, Healthcare Diagnostics, Financial Trust and Compliance, Cybersecurity, Privacy & AI Safety: A Comprehensive Survey, Roadmap & Implementation Blueprint

Kiarash Ahi

专题命中 AI治理与伦理 :safety(title,abstract);AI safety(title);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18328 2025-04-28 cs.CY 86%

AI Safety Assurance for Automated Vehicles: A Survey on Research, Standardization, Regulation

Lars Ullrich, Michael Buchholz, Klaus Dietmayer, Knut Graichen

专题命中 AI治理与伦理 :safety(title,abstract);AI safety(title);分类 cs.CY

Comments Published in IEEE Transactions on Intelligent Vehicles,15 November 2024

Journal ref IEEE Transactions on Intelligent Vehicles 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01225 2025-02-04 cs.CR cs.AI 86%

The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models

Zhiyuan Xu, Joseph Gardiner, Sana Belguith

专题命中 AI治理与伦理 :safety(title,abstract);alignment(title);分类 cs.AI

Comments 12 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14570 2023-11-27 cs.AI physics.med-ph 86%

RAISE -- Radiology AI Safety, an End-to-end lifecycle approach

M. Jorge Cardoso, Julia Moosbauer, Tessa S. Cook, B. Selnur Erdal, Brad Genereaux, Vikash Gupta, Bennett A. Landman, Tiarna Lee, Parashkev Nachev, Elanchezhian Somasundaram, Ronald M. Summers, Khaled Younis, Sebastien Ourselin, Franz MJ Pfister

专题命中 AI治理与伦理 :safety(title,abstract);AI safety(title);分类 cs.AI

Comments 14 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.00430 2019-07-02 cs.AI 86%

Requisite Variety in Ethical Utility Functions for AI Value Alignment

Nadisha-Marie Aliman, Leon Kester

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract,comments);AI safety(abstract,comments);分类 cs.AI

Comments IJCAI 2019 AI Safety Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05748 2025-03-11 cs.CY cs.AI 86%

Alignment, Agency and Autonomy in Frontier AI: A Systems Engineering Perspective

Krti Tallam

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08441 2025-02-14 cs.CL cs.AI 86%

Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation

Riccardo Cantini, Giada Cosenza, Alessio Orsino, Domenico Talia

专题命中 AI治理与伦理 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01042 2024-11-19 cs.LG cs.AI cs.CY 86%

Introduction to AI Safety, Ethics, and Society

Dan Hendrycks

专题命中 AI治理与伦理 :safety(title);AI safety(title);分类 cs.AI、cs.CY、cs.LG

Comments 603 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19621 2026-07-30 cs.LG eess.IV stat.ML 版本更新 85%

AI Alignment in Medical Imaging: Unveiling Hidden Biases Through Counterfactual Analysis

医学影像中的AI对齐:通过反事实分析揭示隐藏偏差

Haroui Ma, Francesco Quinzan, Theresa Willem, Stefan Bauer

机构 * School of Computation, Information and Technology, Technical University of Munich(技术大学慕尼黑计算、信息与技术学院) Department of Computer Science, University of Oxford(牛津大学计算机科学系) Department of Engineering Science, University of Oxford(牛津大学工程科学系) Helmholtz AI, Munich, Germany(海德堡人工智能研究所,德国慕尼黑)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.LG

AI总结 本研究针对医学影像ML系统的偏差问题,提出结合条件潜在扩散模型与统计假设检验的统计框架,在多数据集上验证其优于基线,为医疗AI安全提供可靠工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13142 2026-01-16 cs.AI cs.HC 85%

Can LLMs Understand What We Cannot Say? Measuring Multilevel Alignment Through Abortion Stigma Across Cognitive, Interpersonal, and Structural Levels

LLMs能否理解我们无法表达的内容?通过堕胎污名的多层级对齐进行测量

Anika Sharma, Malavika Mampally, Chidaksh Ravuru, Kandyce Brennan, Neil Gaikwad

机构 * Society-Centered AI Lab(以社会为中心的人工智能实验室) School of Nursing(护理学院)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 研究发现LLMs在认知、人际和结构性层面对堕胎污名的理解不一致,缺乏对多维度心理构造的 coherent 理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04394 2026-01-09 cs.CL 85%

ARREST: Adversarial Resilient Regulation Enhancing Safety and Truth in Large Language Models

ARREST: 对抗鲁棒调节提升大语言模型的安全性与真实性

Sharanya Dasgupta, Arkaprabha Basu, Sujoy Nath, Swagatam Das

机构 * Electronics and Communication Sciences Unit (ECSU), Indian Statistical Institute Kolkata, University of Surrey, and Indian Institute Of Technology Delhi(电子与通信科学单元(ECSU)、印度统计研究所科钦分校、萨里大学和印度理工学院德里)

专题命中 AI治理与伦理 :safety(title,abstract);alignment(abstract);RLHF(abstract);分类 cs.CL

AI总结 ARREST通过对抗鲁棒调节提升大语言模型的安全性与真实性,通过外部网络纠正表征不匹配并生成软拒绝。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12271 2025-11-18 cs.AI 85%

MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning

Zhiyu An, Wan Du

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

Comments Accepted for AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16048 2023-10-25 cs.AI cs.CL cs.CY cs.HC cs.LG 85%

AI Alignment and Social Choice: Fundamental Limitations and Policy Implications

Abhilash Mishra

专题命中 AI治理与伦理 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 10 pages, no figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.03174 2023-03-07 cs.AI cs.CY cs.GT cs.MA econ.GN q-fin.EC 85%

Both eyes open: Vigilant Incentives help Regulatory Markets improve AI Safety

Paolo Bova, Alessandro Di Stefano, The Anh Han

专题命中 AI治理与伦理 :safety(title);AI safety(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20333 2025-08-29 cs.LG cs.AI cs.CL cs.DC 85%

Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs

Md Abdullah Al Mamun, Ihsen Alouani, Nael Abu-Ghazaleh

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00253 2025-06-10 cs.CL cs.AI cs.CY 85%

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

Lihao Sun, Chengzhi Mao, Valentin Hofmann, Xuechunzi Bai

机构 * University of Chicago(芝加哥大学) Rutgers University(罗格斯大学) Allen Institute for AI(人工智能研究所) University of Washington(华盛顿大学)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02153 2025-02-05 cs.AI cs.CL cs.LG 85%

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing

Thien Q. Tran, Akifumi Wachi, Rei Sato, Takumi Tanabe, Youhei Akimoto

专题命中 AI治理与伦理 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12405 2025-01-23 cs.CY cs.AI cs.CL 85%

Scopes of Alignment

Kush R. Varshney, Zahra Ashktorab, Djallel Bouneffouf, Matthew Riemer, Justin D. Weisz

专题命中 AI治理与伦理 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.CY

Comments The 2nd International Workshop on AI Governance (AIGOV) held in conjunction with AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01818 2025-01-17 cs.AI cs.CY cs.LG cs.MA 85%

Hybrid Approaches for Moral Value Alignment in AI Agents: a Manifesto

Elizaveta Tennant, Stephen Hailes, Mirco Musolesi

专题命中 AI治理与伦理 :alignment(title);safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01525 2023-10-23 cs.CV 85%

VisAlign: Dataset for Measuring the Degree of Alignment between AI and Humans in Visual Perception

Jiyoung Lee, Seungho Kim, Seunghyun Won, Joonseok Lee, Marzyeh Ghassemi, James Thorne, Jaeseok Choi, O-Kil Kwon, Edward Choi

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);AI safety(abstract)

Comments Published as a conference paper at NeurIPS 2023 (Track on Datasets and Benchmarks)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02231 2023-06-14 cs.CY cs.AI cs.LG 85%

Connecting the Dots in Trustworthy Artificial Intelligence: From AI Principles, Ethics, and Key Requirements to Responsible AI Systems and Regulation

Natalia Díaz-Rodríguez, Javier Del Ser, Mark Coeckelbergh, Marcos López de Prado, Enrique Herrera-Viedma, Francisco Herrera

专题命中 AI治理与伦理 :trustworthy(title,abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 30 pages, 5 figures, under second review

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09292 2022-02-21 eess.SY cs.AI cs.CY cs.LG cs.SE cs.SY 85%

System Safety and Artificial Intelligence

Roel I. J. Dobbe

专题命中 AI治理与伦理 :safety(title,abstract);AI safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments To appear in: Oxford Handbook on AI Governance (Oxford University Press, 2022 forthcoming)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04289 2024-04-09 cs.AI cs.HC cs.LG 84%

Designing for Human-Agent Alignment: Understanding what humans want from their agents

Nitesh Goyal, Minsuk Chang, Michael Terry

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG

Comments Human-AI Alignment, Human-Agent Alignment, Agents, Generative AI, Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10217 2025-07-02 cs.CY 84%

Data-Centric Safety and Ethical Measures for Data and AI Governance

Srija Chakraborty

专题命中 AI治理与伦理 :safety(title,abstract);red teaming(abstract);分类 cs.CY;AI safety(comments)

Comments Paper accepted and presented at the AAAI 2025 Workshop on Datasets and Evaluators of AI Safety https://sites.google.com/view/datasafe25/home

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09475 2026-06-10 cs.AI cs.LG 版本更新 84%

Emergent alignment and the projectability of ethical personas

涌现对齐与伦理人格的可投射性

Guillermo Del Pinal, Youngchan Lee, Calum McNamara, Alejandro Perez Carballo

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Indiana University Bloomington(印第安纳大学布卢明顿分校)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 研究微调大语言模型在窄任务上如何引发广泛对齐行为,通过宪法AI方法赋予模型伦理人格,发现窄对齐可投射到未训练类别,并提出对齐策略应评估可投射性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08172 2026-06-09 cs.HC cs.AI cs.CY 新提交 84%

The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In

人类与LLM交互的治理:安全门控、文明引导与情感默认锁定

Manuele Reani, Hongjian Zhang, Hongyu Tian

机构 * School of Management and Economics, The Chinese University of Hong Kong, Shenzhen, China(管理学院与经济学学院,香港中文大学(深圳))

专题命中 AI治理与伦理 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究通过确定性多智能体评估流水线,测量LLM在长程对话中的提示可引导性和风格漂移,提出区分安全门控、文明引导和情感默认锁定的治理框架,揭示提供商对交互形式的控制对多元性、自主性和民主能动性的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01346 2026-04-08 cs.CR cs.AI cs.LG cs.RO 84%

Safety, Security, and Cognitive Risks in World Models

世界模型中的安全性、安全性和认知风险

Manoj Parmar

机构 * SovereignAI Security Labs(SovereignAI安全实验室)

专题命中 AI治理与伦理 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了世界模型在自主决策中的安全、安全及认知风险,提出了轨迹持久性和表征风险的定义,并通过实验验证了对抗攻击的效果,强调了对世界模型的严谨性要求。

Comments version 2, 29 pages, 1 figure (6 panels), 3 tables. Empirical proof-of-concept on GRU/RSSM/DreamerV3 architectures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12907 2025-06-17 cs.AI cs.CY cs.GT cs.HC 84%

Roadmap on Incentive Compatibility for AI Alignment and Governance in Sociotechnical Systems

Zhaowei Zhang, Fengshuo Bai, Mingzhi Wang, Haoyang Ye, Chengdong Ma, Yaodong Yang

机构 * Institute for Artificial Intelligence, Peking University(人工智能研究院,北京大学) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Shanghai Jiao Tong University(上海交通大学) Zhongguancun Academy(中关村学院)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15274 2025-06-02 cs.AI cs.LG 84%

From the Pursuit of Universal AGI Architecture to Systematic Approach to Heterogenous AGI: Addressing Alignment, Energy, & AGI Grand Challenges

Eren Kurshan

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG

Comments Categories: Artificial Intelligence; AI; Artificial General Intelligence; AGI; System Design; System Architecture Preprint International Journal on Semantic Computing Vol. 18, No. 03, pp. 465-500

Journal ref International Journal on Semantic Computing Vol. 18, No. 03, pp. 465-500 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏