arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1832 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1832 篇

2101.02032 2021-08-24 cs.CY cs.AI 62%

Socially Responsible AI Algorithms: Issues, Purposes, and Challenges

Lu Cheng, Kush R. Varshney, Huan Liu

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 45 pages, 8 figures

Journal ref Journal of Artificial Intelligence Research 71 (2021) 1137-1181

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06216 2021-08-16 cs.IR cs.AI cs.CL cs.SI 62%

MAIR: Framework for mining relationships between research articles, strategies, and regulations in the field of explainable artificial intelligence

Stanisław Gizinski, Michał Kuzba, Bartosz Pielinski, Julian Sienkiewicz, Stanisław Łaniewski, Przemysław Biecek

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01764 2021-08-05 cs.CL cs.AI 62%

Q-Pain: A Question Answering Dataset to Measure Social Bias in Pain Management

Cécile Logé, Emily Ross, David Yaw Amoah Dadey, Saahil Jain, Adriel Saporta, Andrew Y. Ng, Pranav Rajpurkar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to the 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.10939 2021-07-26 cs.IR cs.CY cs.LG 62%

What are you optimizing for? Aligning Recommender Systems with Human Values

Jonathan Stray, Ivan Vendrov, Jeremy Nixon, Steven Adler, Dylan Hadfield-Menell

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG

Comments Originally presented at the ICML 2020 Participatory Approaches to Machine Learning workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04243 2021-05-12 cs.CV cs.AI cs.CY 62%

Estimating and Improving Fairness with Adversarial Learning

Xiaoxiao Li, Ziteng Cui, Yifan Wu, Lin Gu, Tatsuya Harada

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 12 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.09469 2021-04-20 cs.LG cs.AI cs.HC 62%

Training Value-Aligned Reinforcement Learning Agents Using a Normative Prior

Md Sultan Al Nahian, Spencer Frazier, Brent Harrison, Mark Riedl

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments (Nahian and Frazier contributed equally to this work)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.01685 2021-03-17 cs.AI cs.LG 62%

Agent Incentives: A Causal Perspective

Tom Everitt, Ryan Carey, Eric Langlois, Pedro A Ortega, Shane Legg

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments In Proceedings of the AAAI 2021 Conference. Supersedes arXiv:1902.09980, arXiv:2001.07118

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.10766 2021-03-01 cs.LG cs.AI stat.ML 62%

Teaching the Old Dog New Tricks: Supervised Learning with Constraints

Fabrizio Detassis, Michele Lombardi, Michela Milano

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06058 2020-12-14 cs.CY cs.AI 62%

Next Wave Artificial Intelligence: Robust, Explainable, Adaptable, Ethical, and Accountable

Odest Chadwicke Jenkins, Daniel Lopresti, Melanie Mitchell

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments A Computing Community Consortium (CCC) white paper, 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.00501 2020-03-25 cs.CY cs.AI 62%

The role of artificial intelligence in achieving the Sustainable Development Goals

Ricardo Vinuesa, Hossein Azizpour, Iolanda Leite, Madeline Balaam, Virginia Dignum, Sami Domisch, Anna Felländer, Simone Langhans, Max Tegmark, Francesco Fuso Nerini

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.11990 2020-03-16 cs.LG cs.AI stat.ML 62%

Deontological Ethics By Monotonicity Shape Constraints

Serena Wang, Maya Gupta

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments AISTATS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.12393 2020-01-17 cs.CY cs.AI physics.soc-ph 62%

To regulate or not: a social dynamics analysis of the race for AI supremacy

The Anh Han, Luis Moniz Pereira, Francisco C. Santos, Tom Lenaerts

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.11452 2019-10-28 cs.LG cs.CY stat.ML 62%

Fairness Sample Complexity and the Case for Human Intervention

Ananth Balashankar, Alyssa Lees

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG

Comments Where is the Human? Bridging the Gap Between AI and HCI, CHI Workshop 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.02134 2019-08-07 cs.CY cs.LG cs.SE 62%

Adapting SQuaRE for Quality Assessment of Artificial Intelligence Systems

Hiroshi Kuwajima, Fuyuki Ishikawa

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.06289 2019-05-16 cs.HC cs.CY cs.LG 62%

A Human-Centered Approach to Interactive Machine Learning

Kory W. Mathewson

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

Comments 4 pages, 4th Multidisciplinary Conference on Reinforcement Learning and Decision Making

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.07261 2019-02-08 cs.CY cs.AI 62%

FactSheets: Increasing Trust in AI Services through Supplier's Declarations of Conformity

Matthew Arnold, Rachel K. E. Bellamy, Michael Hind, Stephanie Houde, Sameep Mehta, Aleksandra Mojsilovic, Ravi Nair, Karthikeyan Natesan Ramamurthy, Darrell Reimer, Alexandra Olteanu, David Piorkowski, Jason Tsay, Kush R. Varshney

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.05239 2019-01-09 cs.HC cs.CY cs.LG cs.SE 62%

Improving fairness in machine learning systems: What do industry practitioners need?

Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé, Miro Dudík, Hanna Wallach

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG

Comments To appear in the 2019 ACM CHI Conference on Human Factors in Computing Systems (CHI 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.06126 2018-09-25 cs.AI cs.LG stat.ML 62%

Bootstrapping with Models: Confidence Intervals for Off-Policy Evaluation

Josiah P. Hanna, Peter Stone, Scott Niekum

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments Published in proceedings of the 16th International Conference on Autonomous Agents and Multi-agent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.09093 2017-10-03 cs.CY cs.AI cs.HC 62%

Beyond opening up the black box: Investigating the role of algorithmic systems in Wikipedian organizational culture

R. Stuart Geiger

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 14 pages, typo fixed in v2

Journal ref Big Data & Society 4(2). 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01451 2026-05-05 cs.CL 61%

Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models

对基于AI的紧急警务调度中的种族偏见进行审计:对十一种大型语言模型的跨语言评估

William Guey, Wei Zhang, Pierrick Bougault, Yi Wang, Bertan Ucar, Vitor D. de Moura, José O. Gomes

机构 * Department of Industrial Engineering, Tsinghua University(清华大学工业工程系) School of Social Sciences, Tsinghua University(清华大学社会科学部) Department of Industrial Engineering, Federal University of Rio de Janeiro(里约热内卢联邦大学工业工程系)

专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.CL

AI总结 本文通过跨语言框架评估11种模型,在19800个输出中发现当事件严重性模糊时种族偏见系统性出现,但当操作优先级由通话内容确定时偏见消失。偏见程度因种族轴而异,宗教外观影响最大,性别次之,种族最小。语言间偏见转移不一致,性别偏见在中文中放大,种族偏见在英文中更明显。

Comments 26 pages, 7 figures. Submitted to Humanities and Social Sciences Communications (Nature) collection on Artificial Intelligence and Emerging Technologies in Public Safety. Code and data: https://github.com/williamguey/llmdispatchbias

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11839 2026-04-15 cs.CR cs.AI 61%

Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents

超越静态沙箱:为自主AI代理的学得能力治理

Bronislav Sidik, Lior Rokach

机构 * Institute for Applied AI Research(应用人工智能研究所) Faculty of Computer and Information Science(计算机与信息科学学院) Ben-Gurion University of the Negev(贝内-约尔大学)

专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.AI

AI总结 本文提出Aethelgard框架,通过学得策略实现AI代理的最小必要能力集,解决能力过度配置问题。

Comments 17 pages (9 content pages), 2 figures, 7 tables. Submitted to NeurIPS 2026 Agent Safety Workshop. Code and dataset available at https://github.com/sidikbro/aethelgard

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22037 2025-12-05 cs.CY 61%

What AI Speaks for Your Community: Polling AI Agents for Public Opinion on Data Center Projects

人工智能为你的社区发声:通过AI代理收集数据中心项目公众意见

Zhifeng Wu, Yuelin Han, Shaolei Ren

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY;trustworthy(comments)

AI总结 本文提出AI代理调查框架,利用大型语言模型评估社区对数据中心项目的意见,以指导负责任的AI发展。

Comments 35 Pages. Accepted to NeurIPS 2025 Workshop on Socially Responsible and Trustworthy Foundation Models (ResponsibleFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09222 2025-08-22 cs.CY 61%

Democratic AI is Possible. The Democracy Levels Framework Shows How It Might Work

Aviv Ovadya, Kyle Redman, Luke Thorburn, Quan Ze Chen, Oliver Smith, Flynn Devine, Andrew Konya, Smitha Milli, Manon Revel, K. J. Kevin Feng, Amy X. Zhang, Bilva Chandra, Michiel A. Bakker, Atoosa Kasirzadeh

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CY

Comments 31 pages. Accepted to the position paper track at ICML 2025. A previous version was presented at the Pluralistic Alignment Workshop at NeurIPS 2024. For ongoing work, see: https://democracylevels.org

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05617 2025-05-20 cs.LG cs.AI cs.CL cs.CV cs.CY 61%

Debiasing Methods for Fairer Neural Models in Vision and Language Research: A Survey

Otávio Parraga, Martin D. More, Christian M. Oliveira, Nathan S. Gavenski, Lucas S. Kupssinskü, Adilson Medronha, Luis V. Moura, Gabriel S. Simões, Rodrigo C. Barros

机构 * Machine Learning Theory and Applications (MALTA) Lab, PUCRS(机器学习理论与应用(MALTA)实验室,PUCRS)

专题命中 AI治理与伦理 :分类 cs.CL、cs.AI、cs.CY;trustworthy(comments)

Comments Submitted to ACM Computing Surveys - Special Issue on Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11579 2025-01-22 cs.CL 61%

HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection

Theo King, Zekun Wu, Adriano Koshiyama, Emre Kazim, Philip Treleaven

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL;safety(comments)

Comments NeurIPS 2024 SoLaR Workshop and NeurIPS 2024 Safety Gen AI Workshop

Journal ref NeurIPS Safe Generative AI Workshop 2024; Workshop on Socially Responsible Language Modelling Research 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09362 2024-10-02 cs.LG 61%

Long-Term Fairness with Unknown Dynamics

Tongxin Yin, Reilly Raab, Mingyan Liu, Yang Liu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG;trustworthy(comments)

Comments Best paper runner-up at ICLR 2023 Workshop on Trustworthy and Reliable Large-Scale Machine Learning Models (Non Archival)

Journal ref Advances in Neural Information Processing Systems 36 (NeurIPS 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09209 2023-07-19 cs.CL cs.AI cs.CY cs.LG 61%

Automated Ableism: An Exploration of Explicit Disability Biases in Sentiment and Toxicity Analysis Models

Pranav Narayanan Venkit, Mukund Srinath, Shomir Wilson

专题命中 AI治理与伦理 :分类 cs.CL、cs.AI、cs.CY;trustworthy(journal_ref)

Comments TrustNLP at ACL 2023

Journal ref Proceedings at The Third Workshop on Trustworthy Natural Language Processing collocated at the 61st Annual Meeting Of The Association For Computational Linguistics. 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00168 2022-11-02 cs.CV cs.LG 61%

Improving Fairness in Image Classification via Sketching

Ruichen Yao, Ziteng Cui, Xiaoxiao Li, Lin Gu

专题命中 AI治理与伦理 :trustworthy(abstract,comments);分类 cs.LG

Comments 8 pages, 2 figures. To appear in 2022 Trustworthy and Socially Responsible Machine Learning (TSRML 2022) co-located with NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16394 2026-08-18 cs.AI cs.IR 新提交 57%

Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152

在块内思考:使用大语言模型生成符合法规的场景的RegulaRAG——以联合国第152号法规为例

Vahid Zolfaghari, Nenad Petrovic, AndrÉ Schamschurko, Alois Knoll

机构 * Technical University of Munich(慕尼黑工业大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 针对LLMs难以结合冗长分层标准的问题,提出RegulaRAG流水线,经实验其在UN R152数据集上元分数最高且鲁棒性强,优于基线系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15354 2026-08-18 cs.AI 新提交 57%

Incoherent by Design? On the Moral Self-Consistency of LLMs

天生不连贯?大型语言模型的道德自我一致性研究

Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum

机构 * Cornell University(康奈尔大学) Cornell Tech(康奈尔科技学院) Nanyang Technological University(南洋理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本研究针对GPT、Mistral、Llama等LLM,在义务论等三大伦理框架下,发现其道德推理存在最高78%的矛盾率,内部不连贯是AI对齐的必要前提。

Comments 88 pages; pages 16 to 88 are the appendix

详情

展开后加载摘要…

URL PDF HTML 收藏