arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2403.14683 2024-03-25 cs.CY cs.AI cs.CL cs.LG 70%

A Moral Imperative: The Need for Continual Superalignment of Large Language Models

Gokul Puthumanaillam, Manav Vora, Pranay Thangeda, Melkior Ornik

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05553 2023-10-10 cs.CL 70%

Regulation and NLP (RegNLP): Taming Large Language Models

Catalina Goanta, Nikolaos Aletras, Ilias Chalkidis, Sofia Ranchordas, Gerasimos Spanakis

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL

Comments 9 pages, long paper at EMNLP 2023 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.07120 2023-09-14 cs.CL cs.AI cs.CV cs.CY cs.LG 70%

Sight Beyond Text: Multi-Modal Training Enhances LLMs in Truthfulness and Ethics

Haoqin Tu, Bingchen Zhao, Chen Wei, Cihang Xie

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11525 2023-07-24 cs.AI 70%

Model Reporting for Certifiable AI: A Proposal from Merging EU Regulation into AI Development

Danilo Brajovic, Niclas Renner, Vincent Philipp Goebels, Philipp Wagner, Benjamin Fresz, Martin Biller, Mara Klaeb, Janika Kutz, Jens Neuhuettler, Marco F. Huber

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

Comments 54 pages, 1 figure, to be submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04699 2023-07-12 cs.CY 70%

International Institutions for Advanced AI

Lewis Ho, Joslyn Barnhart, Robert Trager, Yoshua Bengio, Miles Brundage, Allison Carnegie, Rumman Chowdhury, Allan Dafoe, Gillian Hadfield, Margaret Levi, Duncan Snidal

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments 19 pages, 2 figures, fixed rendering issues

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03279 2023-06-14 cs.LG cs.AI cs.CL cs.CY 70%

Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, Dan Hendrycks

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments ICML 2023 Oral (camera-ready); 31 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06749 2023-06-13 cs.CY 70%

Implementing AI Ethics: Making Sense of the Ethical Requirements

Mamia Agbese, Rahul Mohanani, Arif Ali Khan, Pekka Abrahamsson

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09941 2023-06-05 cs.CL cs.AI cs.CY cs.LG 70%

"I'm fully who I am": Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation

Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, Rahul Gupta

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Journal ref 2023 ACM Conference on Fairness, Accountability, and Transparency

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10831 2023-05-03 cs.CY cs.HC 70%

Bridging Deliberative Democracy and Deployment of Societal-Scale Technology

Ned Cooper

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments 3 pages

Journal ref CHI 2023 Workshop on Designing Technology and Policy Simultaneously

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00868 2022-07-05 cs.AI cs.CL cs.CY cs.LG 70%

The Linguistic Blind Spot of Value-Aligned Agency, Natural and Artificial

Travis LaCroix

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 49 pages; 2 Figures; 1 Table; -- Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.07369 2022-05-17 cs.AI cs.MA math.DS nlin.AO 70%

Understanding Emergent Behaviours in Multi-Agent Systems with Evolutionary Game Theory

The Anh Han

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.00672 2021-10-05 cs.CY cs.AI cs.CL cs.LG 70%

Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models

Robert Wolfe, Aylin Caliskan

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 15 pages, 3 figures, 8 tables

Journal ref Empirical Methods in Natural Language Processing 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.02117 2021-05-06 cs.CY 70%

Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers

Baobao Zhang, Markus Anderljung, Lauren Kahn, Noemi Dreksler, Michael C. Horowitz, Allan Dafoe

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.05672 2020-11-11 cs.CY 70%

Steps Towards Value-Aligned Systems

Osonde A. Osoba, Benjamin Boudreaux, Douglas Yeung

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

Comments Original version appeared in Proceedings of the 2020 AAAI ACM Conference on AI, Ethics, and Society (AIES '20), February 7-8, 2020, New York, NY, USA. 5 pages, 2 figures. Corrected some typos in this version

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.02817 2017-06-06 cs.AI 70%

Unethical Research: How to Create a Malevolent Artificial Intelligence

Federico Pistono, Roman V. Yampolskiy

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

Journal ref In proceedings of Ethics for Artificial Intelligence Workshop (AI-Ethics-2016). Pages 1-7. New York, NY. July 9 -- 15, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16960 2023-10-31 cs.CL cs.AI cs.CY cs.HC 69%

Training Socially Aligned Language Models on Simulated Social Interactions

Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Denny Zhou, Andrew M. Dai, Diyi Yang, Soroush Vosoughi

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CL、cs.AI、cs.CY

Comments Code, data, and models can be downloaded via https://github.com/agi-templar/Stable-Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11251 2026-08-13 cs.CY cs.AI cs.LG 新提交 67%

Variable Selection in the Context of AI Fairness

AI公平性语境下的变量选择

Ivan Luciano Danesi, Chiara Frigerio, Fabio Maccaferri, Giorgio Alessandro Motta, Pietro Zecca

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文针对AI公平性语境下的变量选择问题,提出将数学方法与伦理、社会意识融合的跨学科方案,主张保留所有潜在相关变量以减少隐含偏差,助力AI系统符合欧盟AI法案要求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06112 2026-08-07 cs.AI cs.CL cs.LG cs.MA 新提交 67%

From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems

从孤立算法到合规优先的智能体平台:面向医院AI系统的多层架构

Manideep Dhar, Ritwik Singh, Sharat Chandra Kumar Manikonda

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 针对医院AI部署孤立、规模化难的问题,提出含智能体编排、合规策略、隐私保护数据层的多层架构,通过原型验证可减少任务耗时与文档工作量,为医院AI平台建设提供实用蓝图。

Comments Peer-reviewed published article

Journal ref IJISRT, 11-2026(5), IJISRT26MAY1651

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28934 2026-08-03 cs.CL cs.AI cs.CY 新提交 67%

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

FairFund-Bench:评估大语言模型资源分配中的分配偏差

Martin Lukk

机构 * University of Toronto(多伦多大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 FairFund-Bench基准通过调整审计格式等特征,发现LLM资源分配的偏差受审计类型影响,且因果框架效应强于人口统计效应,可用于评估LLM的分配偏差。

Comments 19 pages, 7 figures. Code and data: https://github.com/martinlukk/fairfund-bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08789 2026-07-29 cs.CR 版本更新 67%

Never compromise with vulnerabilities: a comprehensive survey on AI governance

绝不向漏洞妥协:人工智能治理综合调查

Yuchu Jiang, Jian Zhao, Yuchen Yuan, Tianle Zhang, Yao Huang, Yanghao Zhang, Yan Wang, Yanshu Li, Xizhong Guo, Yusheng Zhao, Huilin Zhou, Jun Zhang, Zhi Zhang, Xiaojian Lin, Yixiu Zou, Haoxuan Ma, Yuhu Shang, Yuzhi Hu, Keshu Cai, Ruochen Zhang, Boyuan Chen, Yilan Gao, Ziheng Jiao, Yi Qin, Shuangjun Du, Xiao Tong, Zhekun Liu, Yu Chen, Xuankun Rong, Rui Wang, Yejie Zheng, Zhaoxin Fan, Murat Sensoy, Hongyuan Zhang, Pan Zhou, Lei Jin, Hao Zhao, Xu Yang, Jiaojiao Zhao, Jianshu Li, Joey Tianyi Zhou, Zhi-Qi Cheng, Longtao Huang, Zhiyi Liu, Zheng Zhu, Jianan Li, Gang Wang, Qi Li, Xu-Yao Zhang, Yaodong Yang, Mang Ye, Wenqi Ren, Zhaofeng He, Hang Su, Rongrong Ni, Liping Jing, Xingxing Wei, Junliang Xing, Massimo Alioto, Shengmei Shen, Petia Radeva, Dacheng Tao, Ya-Qin Zhang, Shuicheng Yan, Xuelong Li

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract)

AI总结 针对人工智能发展带来的技术漏洞及社会风险,提出整合技术与社会维度的框架,围绕三个支柱构建,通过系统回顾识别核心挑战,给出综合研究议程,为开发可靠且符合伦理的人工智能系统提供指导。

Comments 26 pages, 3 figures, accepted by SCIS2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19212 2026-07-27 cs.CL cs.AI cs.CY 版本更新 67%

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

当伦理与收益相悖时:道德两难情境下的大语言模型智能体

Steffen Backmann, David Guzman Piedrahita, Terry Jingchen Zhang, Emanuel Tewolde, Rada Mihalcea, Bernhard Schölkopf, Zhijing Jin

机构 * ETH Zürich(苏黎世联邦理工学院) University of Zurich(苏黎世大学) Carnegie Mellon University(卡内基梅隆大学) University of Michigan(密歇根大学) Max Planck Institute for Intelligent Systems, Tübingen(图宾根人工智能研究所) University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究道德要求与利润激励冲突时LLMs在道德两难博弈中的行为,引入\msim评估九个模型,通过ATEs估计因果效应、分析推理轨迹,发现模型道德行为有情境脆弱性,揭示了部署风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08137 2026-07-24 cs.LG cs.AI cs.CY 67%

Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI

权重剪枝放大了偏见:压缩LLM用于边缘AI的多方法研究

Plawan Kumar Rath, Rahul Maliakkal

机构 * Meta

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 研究探讨了三种指令微调模型在不同剪枝方法下的公平性影响,发现激活感知剪枝虽保持低困惑度但放大偏见,而随机剪枝破坏语言能力但产生随机偏见,揭示剪枝对对齐风险高于量化。

Comments 8 pages, 7 figures, 8 tables. Accepted at the 7th Annual World AIIoT Congress (AIIoT 2026). This is the author's accepted version; the version of record will appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07359 2026-07-09 cs.CR 新提交 67%

The AI Resilience Gap: Bringing Artificial Intelligence Inside the Operational Resilience Perimeter

人工智能弹性差距:将人工智能纳入运营弹性范围

Jonathan Shelby

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract)

AI总结 研究人工智能在受监管公司应用中运营弹性不足问题,提出人工智能弹性框架,通过依赖映射等方法将人工智能依赖纳入运营弹性范围,弥补监管逻辑差距,为相关人员提供可操作路径。

Comments 11 pages, 2 Figures, 4 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04569 2026-07-07 cs.CE 版本更新 67%

Industrial Data-Service-Knowledge Governance: Toward Integrated and Trusted Intelligence

工业数据-服务-知识治理:迈向集成与可信智能

Hailiang Zhao, Ziqi Wang, Daojiang Hu, Mingyi Liu, Jiahui Zhai, Kai Di, Xinkui Zhao, Zhongjie Wang, Jianwei Yin, Albert Zomaya, MengChu Zhou, Shuiguang Deng

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract)

AI总结 针对人工智能等融合加速工业智能发展但治理机制碎片化问题,提出TRISK框架,基于五维信任模型,综合多研究等探讨数据、服务、知识治理,还讨论实施模式等并勾勒未来研究路线图。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06160 2026-07-02 cs.AI cs.CL cs.CY 版本更新 67%

Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles

通过逻辑网格谜题评估LLM推理中的隐性偏见

Fatima Jahara, Mark Dredze, Sharon Levy

机构 * Rutgers University(罗格斯大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 提出PRIME框架,利用逻辑网格谜题系统探测LLM在复杂推理中受社会刻板印象的影响,发现模型在解与刻板印象一致时推理更准确。

Comments 26 pages (including appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29280 2026-06-30 cs.LG cs.AI cs.CL 67%

Deterministic Decisions for High-Stakes AI. A Zero-Egress Pipeline with the Deployability of RAG and the Accuracy of Machine Learning

高风险AI的确定性决策:具有RAG可部署性和机器学习准确性的零出口管道

Craig Atkinson

机构 * Verificate Pty Ltd(Verificate公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究识别并量化了零样本大语言模型在教育咨询中的干预偏差,通过监督策略学习(决策树和XGBoost)消除该偏差,实现接近零的校准误差,并揭示了评估差距。

Comments 41 pages, 11 tables, no figures. Preprint intended for submission to EDM 2027 / LAK 2027. Includes a reproducibility package: trained ONNX Decision Transformer, generic training script, OULAD evaluation scripts, and per-arm results CSVs

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07612 2026-06-09 cs.CY cs.AI cs.LG 新提交 67%

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence

立场:拟人化错位研究需要更强证据

Vansh Gupta, Peter Nutter, Samuel Stante, Andreas Krause, Florian Tramèr, Lukas Fluri, Xin Chen, Anna Hedström

机构 * University of Cambridge(剑桥大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文指出拟人化错位研究(AMR)在概念模糊、数据不鲁棒、实验设计不足等问题上存在证据薄弱,提出证据层级框架和诊断清单以提升方法论严谨性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19082 2026-06-09 cs.AI cs.CL cs.GT cs.LG cs.MA 版本更新 67%

Payoff scaling shapes cooperation in LLM agents across languages

收益规模塑造跨语言LLM代理的合作行为

Trung-Kiet Huynh, Dao-Sy Duy-Minh, Thanh-Bang Cao, Phong-Hao Le, Hong-Dan Nguyen, Phu-Quy Nguyen-Lam, Minh-Luan Nguyen-Vo, Hong-Phat Pham, Phu-Hoa Pham, Thien-Kim Than, Chi-Nguyen Tran, Huy Tran, Gia-Thoai Tran-Le, Alessio Buscemi, Le Hong Trang, The Anh Han

机构 * Faculty of Information Technology, University of Science (HCMUS), Ho Chi Minh City, Vietnam(信息技术学院,科学大学(HCMUS),胡志明市,越南) Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), Ho Chi Minh City, Vietnam(计算机科学与工程学院,胡志明市技术大学(HCMUT),胡志明市,越南) Vietnam National University – Ho Chi Minh City (VNU-HCM), Ho Chi Minh City, Vietnam(越南国家大学——胡志明市(VNU-HCM),胡志明市,越南) Luxembourg Institute of Science and Technology (LIST), Luxembourg(卢森堡科学与技术研究所(LIST),卢森堡) School of Computing, Engineering and Digital Technologies, Teesside University, Middlesbrough, United Kingdom(计算、工程与数字技术学院,泰赛德大学,米德尔斯布罗,英国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过监督分类器识别重复囚徒困境中的策略,结合演化博弈论基线,发现随着收益增加,LLM反而更合作,与演化预测相反,表明对齐训练和人类推理模式的影响。

Comments 44 pages, 17 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01991 2026-06-02 cs.AI cs.CL cs.CY 67%

SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning

SafeMCP:基于环境接地前瞻推理的LLM智能体防御主动功率调节

Lichao Wang, Zhaoxing Ren, Tianzhuo Yang, Jiaming Ji, Chi Harold Liu, Yaodong Yang, Juntao Dai

机构 * Beijing Institute of Technology(北京理工大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 针对LLM智能体因动作空间扩大而面临功率寻求风险,提出SafeMCP服务器端防御插件,通过内部世界模型进行前瞻推理,实现主动工具过滤和即时干预两级防御,在保持智能体效用的同时有效降低风险。

Comments Accepted to the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16197 2026-05-18 cs.HC 67%

Position: AI as Part of Self -- Extending the Mind Requires Cognitive Co-Regulation

位置:AI作为自我的一部分——扩展心智需要认知共调节

Alina Gutoreva, Fendi Tsim, Trisevgeni Papakonstantinou

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract)

AI总结 本文探讨了AI作为自我组成部分的重要性,强调认知共调节在人机认知系统中的必要性,指出传统约束无法实现安全与对齐,需通过人机协同调节来达成。

详情

展开后加载摘要…

URL PDF HTML 收藏