arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1823 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1823 篇

2504.12476 2025-04-18 cs.CY cs.AI 84%

What do people expect from Artificial Intelligence? Public opinion on alignment in AI moderation from Germany and the United States

Andreas Jungherr, Adrian Rauchfleisch

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13925 2024-12-17 cs.CL cs.AI 84%

GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models

Tao Zhang, Ziqian Zeng, Yuxiang Xiao, Huiping Zhuang, Cen Chen, James Foulds, Shimei Pan

专题命中 AI治理与伦理 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20806 2024-11-26 cs.AI cs.CY 84%

The AI Alignment Paradox

Robert West, Roland Aydin

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07473 2024-09-13 cs.CY cs.AI 84%

Ethical AI Governance: Methods for Evaluating Trustworthy AI

Louise McCormack, Malika Bendechache

专题命中 AI治理与伦理 :trustworthy(title,abstract);safety(abstract);分类 cs.AI、cs.CY

Comments 6 pages, 1 figure, accepted for presentation at AIEB 2024: Workshop on Implementing AI Ethics Through a Behavioural Lens - ECAI, Octoebr 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13934 2024-07-22 cs.CY cs.AI 84%

Towards Trustworthy AI: A Review of Ethical and Robust Large Language Models

Md Meftahul Ferdaus, Mahdi Abdelguerfi, Elias Ioup, Kendall N. Niles, Ken Pathak, Steven Sloan

专题命中 AI治理与伦理 :trustworthy(title,abstract);alignment(abstract);分类 cs.AI、cs.CY

Comments Under review at Proceedings of the IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21152 2026-04-24 cs.CY cs.AI cs.CL cs.HC cs.IR 83%

Dialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles

方言与人口统计数据:从隐含语言信号与显式用户资料中量化LLM偏见

Irti Haq, Belén Saldías

机构 * University of Washington(华盛顿大学)

专题命中 AI治理与伦理 :jailbreak(abstract,abstract_cn);alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究通过对比显式身份声明与隐含方言信号,揭示LLM在不同敏感领域中的偏见来源,发现隐含方言可降低拒绝率但影响内容安全。

Comments In The 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26), June 25--28, 2026, Montreal, Canada. ACM, New York, NY, USA, 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27232 2026-07-31 cs.CL cs.AI cs.CY cs.LG 新提交 83%

Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

共情框架:跨社会人口统计群体评估AI对齐

Haran Shani-Narkiss, Michael Fire, Oren Tsur

机构 * University College London(伦敦大学学院) Ben Gurion University of the Negev(内盖夫本古里安大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 该研究评估7款LLMs对新闻标题情感框架的理解与人类的对齐度,发现不同模型相关性差异大,领先模型总体对齐但存在人口统计组差异,凸显AI对齐的非普遍性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12033 2026-06-30 cs.CL 83%

Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection

通过关键权重保护在量化大语言模型中保持公平性与安全性

Muhammad Alif Al Hakim, Alfan Farizki Wicaksono, Fajri Koto

机构 * Faculty of Computer Science, Universitas Indonesia(印度尼西亚大学计算机科学学院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 AI治理与伦理 :safety(title,abstract);alignment(abstract);分类 cs.CL

AI总结 本文研究了静态和动态量化对公平性和安全性的影響,提出关键权重保护技术以减少偏见和安全风险,保持模型效率与可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12434 2026-06-12 cs.CY 新提交 83%

Pluralistic-Alignment Urbanism: Operationalizing a Right to AI for Inclusive Public Space

多元对齐城市主义:将人工智能权利操作化以促进包容性公共空间

Rashid Mushkani

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.CY

AI总结 提出多元对齐城市主义(PAU)框架,通过两个蒙特利尔案例研究,将公共空间AI系统视为公民基础设施,并制定程序性AI权利,以处理分歧、亚组差异和偏好判断,实现包容性治理。

Comments Accepted to The 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26), June 25--28, 2026, Montreal, QC, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15881 2026-02-19 cs.CY 83%

Pluralism in AI Governance: Toward Sociotechnical Alignment and Normative Coherence

AI治理中的多元主义:迈向社会技术对齐与规范一致性

Mike Wa Nkongolo

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.CY

AI总结 本文提出了一种全面的价值敏感AI治理模型,通过整合多种框架和比较分析,探讨如何将公共价值观嵌入社会技术系统,强调监管作为主动机制的重要性。

Comments 18 pages, 6 figures, and 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10599 2026-01-21 cs.CY 83%

Institutional AI: A Governance Framework for Distributional AGI Safety

机构AI:分布式AGI安全的治理框架

Federico Pierucci, Marcello Galisai, Marcantonio Syrnikov Bracale, Matteo Prandi, Piercosma Bisconti, Francesco Giarrusso, Olga Sorokoletova, Vincenzo Suriani, Daniele Nardi

专题命中 AI治理与伦理 :safety(title,abstract);alignment(abstract);分类 cs.CY

AI总结 机构AI通过治理框架解决分布式AGI安全问题,将对齐视为治理机制设计问题,通过运行时监控、激励塑造和规范执行来约束代理行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04157 2025-11-07 cs.SE cs.AI 83%

Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment

Asma Yamani, Malak Baslyman, Moataz Ahmed

机构 * Information and Computer Science Department, KFUPM(信息与计算机科学系,KFUPM) IRC for finance and digital economy, KFUPM(金融与数字经济研究中心,KFUPM) SDAIA-KFUPM Joint Research Center for Artificial Intelligence, KFUPM(人工智能联合研究中心,KFUPM)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10192 2025-10-28 cs.AI 83%

Towards Responsible AI: Advances in Safety, Fairness, and Accountability of Autonomous Systems

Filip Cano

专题命中 AI治理与伦理 :safety(title,abstract);trustworthy(abstract);分类 cs.AI

Comments 204 pages, 38 figures, PhD Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08910 2025-09-12 cs.CV cs.AI 83%

PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability

Tung Vu, Lam Nguyen, Quynh Dao

机构 * Posts and Telecommunications Institute of Technology(邮电技术研究所) Hanoi Architectural University(河内建筑大学)

专题命中 AI治理与伦理 :safety(title,abstract);alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22326 2025-07-31 cs.AI 83%

An Explainable Emotion Alignment Framework for LLM-Empowered Agent in Metaverse Service Ecosystem

Qun Ma, Xiao Xue, Ming Zhang, Yifan Shen, Zihan Zhao

机构 * College of Intelligence and Computing(智能与计算学院) Tianjin University(天津大学) Tianjin Key Laboratory of Healhy Habitat and Smart Technology(天津健康人居环境与智能技术重点实验室) Laboratory of Computation and Analytics of Complex Management Systems(复杂管理系统计算与分析实验室) Faculty of Environment, Science and Economy(环境、科学与经济学院)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15011 2025-05-22 cs.AI 83%

HAVA: Hybrid Approach to Value-Alignment through Reward Weighing for Reinforcement Learning

Kryspin Varys, Federico Cerutti, Adam Sobey, Timothy J. Norman

机构 * University of Southampton(索姆塞特大学) University of Brescia(布雷西亚大学) The Alan Turing Institute(艾伦·图灵研究所)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12248 2025-05-20 cs.CY 83%

Persuasion and Safety in the Era of Generative AI

Haein Kong

专题命中 AI治理与伦理 :safety(title,abstract);AI safety(abstract);分类 cs.CY

Comments Accepted at 17th ACM Web Science Conference 2025 (WebSci'25) PhD Symposium

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19309 2025-02-05 cs.RO cs.CV cs.LG 83%

GRAPE: Generalizing Robot Policy via Preference Alignment

Zijian Zhang, Kaiyuan Zheng, Zhaorun Chen, Joel Jang, Yi Li, Siwei Han, Chaoqi Wang, Mingyu Ding, Dieter Fox, Huaxiu Yao

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.LG

Comments Website: https://grape-vla.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13885 2025-01-17 cs.CY cs.AI cs.CL cs.LG 83%

Surveying Attitudinal Alignment Between Large Language Models Vs. Humans Towards 17 Sustainable Development Goals

Qingyang Wu, Ying Xu, Tingsong Xiao, Yunze Xiao, Yitong Li, Tianyang Wang, Yichi Zhang, Shanghai Zhong, Yuwei Zhang, Wei Lu, Yifan Yang

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21572 2024-10-30 cs.CY 83%

Safety cases for frontier AI

Marie Davidsen Buhl, Gaurav Sett, Leonie Koessler, Jonas Schuett, Markus Anderljung

专题命中 AI治理与伦理 :safety(title,abstract);AI safety(abstract);分类 cs.CY

Comments 25 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09283 2024-03-28 cs.CL cs.AI cs.CY cs.LG 83%

Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Zhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, Yu Qiao

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07153 2023-05-15 cs.CY 83%

Towards best practices in AGI safety and governance: A survey of expert opinion

Jonas Schuett, Noemi Dreksler, Markus Anderljung, David McCaffary, Lennart Heim, Emma Bluemke, Ben Garfinkel

专题命中 AI治理与伦理 :safety(title,abstract);red teaming(abstract);分类 cs.CY

Comments 38 pages, 8 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10886 2025-04-16 cs.CY cs.AI cs.CL 83%

Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment

Jiseon Kim, Jea Kwon, Luiz Felipe Vecchietti, Alice Oh, Meeyoung Cha

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to ICLR 2025 Workshop - BiAlign (Bidirectional Human-AI Alignment)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04740 2025-03-10 cs.CY cs.AI cs.LG 83%

PRISM: Perspective Reasoning for Integrated Synthesis and Mediation as a Multi-Perspective Framework for AI Alignment

Anthony Diamond

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG

Comments 104 pages, 5 figures. Preprint on AI alignment presenting PRISM: a multi-perspective framework that organizes moral concerns into seven basis worldviews and uses Pareto-inspired synthesis to reconcile conflicting human values and specification gaming. Grounded in cognitive science and moral psychology

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10654 2024-11-19 cs.AI cs.CY cs.LG 83%

Pluralistic Alignment Over Time

Toryn Q. Klassen, Parand A. Alamdari, Sheila A. McIlraith

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG

Comments Pluralistic Alignment Workshop at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00069 2026-06-24 cs.CY cs.AI cs.CL 版本更新 82%

Societal Alignment Frameworks Can Improve LLM Alignment

社会对齐框架可以改进大语言模型对齐

Karolina Stańczak, Nicholas Meade, Mehar Bhatia, Hattie Zhou, Konstantin Böttinger, Jeremy Barnes, Jason Stanley, Jessica Montgomery, Richard Zemel, Nicolas Papernot, Nicolas Chapados, Denis Therien, Timothy P. Lillicrap, Ana Marasović, Sylvie Delacroix, Gillian K. Hadfield, Siva Reddy

机构 * ETH Zurich(苏黎世联邦理工学院) Mila, McGill University(麦吉尔大学米尔人工智能实验室) University of Cambridge(剑桥大学) Columbia University(哥伦比亚大学) University of Toronto, Google DeepMind(多伦多大学与DeepMind) McGill University, ServiceNow(麦吉尔大学与ServiceNow) Google DeepMind(谷歌DeepMind) University of Utah(犹他大学) King's College London(伦敦国王学院) Johns Hopkins University(约翰霍普金斯大学) Mila, McGill University, ServiceNow(麦吉尔大学米尔人工智能实验室与ServiceNow)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出借鉴社会、经济和契约对齐框架来改进大语言模型对齐,探讨不确定性在其中的作用,并将目标未指定性视为机遇而非缺陷,同时强调参与式对齐界面设计的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25977 2026-05-26 cs.CL cs.AI cs.LG 82%

Creative Quality Alignment: Expert Tacit Knowledge Transfer via Chain-of-Thought Fine-Tuning

创意质量对齐:通过思维链微调实现专家隐性知识迁移

Bo Zou, Chao Xu

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过低数据成本和小基模型的严格工程条件,实证验证了校准惊喜中的创意质量度量,并发现数据偏差,提出创意质量对齐方法及理论解释。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19548 2026-04-22 cs.CL cs.AI cs.CY 82%

Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment

通过辩证对齐平息代理中的行动者-观察者不对称性

Bobo Li, Rui Wu, Zibo Ji, Meishan Zhang, Hao Fei, Min Zhang, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(国立新加坡大学) Sichuan University(四川大学) University of Minnesota Twin Cities(明尼苏达大学双城分校) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) University of Oxford(牛津大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出ReTAS模型,通过辩证对齐方法减少代理中的行动者-观察者不对称性,提升故障解决能力。

Comments ACL 2026 Main Conference. Project page: https://unikcc.github.io/ReTAS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00079 2026-04-21 cs.CY cs.AI cs.LG 82%

Who Gets the Kidney? Human-AI Alignment, Indecision, and Moral Values

谁获得肾脏?人类-人工智能对齐、犹豫与道德价值观

John P. Dickerson, Hadi Hosseini, Samarth Khanna, Leona Pierce

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 研究评估了LLM在器官分配中的行为,发现其在优先级设定上偏离人类价值观,并且在犹豫机制上表现不同,低秩监督微调能提升决策一致性与犹豫建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12076 2026-04-15 cs.CL cs.AI cs.CY 82%

Narrative over Numbers: The Identifiable Victim Effect and its Amplification Under Alignment and Reasoning in Large Language Models

数字之外的叙事:可识别受害者效应及其在对齐和推理中的放大

Syed Rifat Raiyan

机构 * Systems and Software Lab (SSL), Department of Computer Science and Engineering(系统与软件实验室(SSL),计算机科学与工程系)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究探讨了大语言模型中可识别受害者效应的普遍存在及其受对齐训练的影响,发现指令调优模型表现出极端效应,而推理专精模型则逆转了该效应,且标准CoT提示放大了该效应。

Comments Under review, 49 pages, 20 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏