arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2601.19768 2026-05-01 cs.AI cs.CR cs.LG 81%

GAVEL: Towards Rule-Based Safety Through Activation Monitoring

GAVEL:通过激活监控实现基于规则的安全性

Shir Rozenfeld, Rahul Pankajakshan, Itay Zloczower, Eyal Lenga, Gilad Gressel, Yisroel Mirsky

机构 * Ben Gurion University of the Negev(本·古里安大学) Amrita Vishwa Vidyapeetham(阿米塔大学)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于规则的激活安全性方法,通过细粒度可解释的激活元素捕捉领域特定行为,提升精度并支持定制化,为可扩展的AI治理奠定基础。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19752 2026-04-23 cs.MA cs.AI cs.CY 81%

Soft-Label Governance for Distributional Safety in Multi-Agent Systems

多智能体系统中的分布安全软标签治理

Aizierjiang Aiersilan, Raeli Savitt

机构 * The George Washington University(乔治·华盛顿大学) SWARM AI Safety(SWARM AI安全)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出SWARM框架,通过连续概率标签提升多智能体系统安全性和治理效果,展示软指标在检测代理游戏中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20805 2026-04-23 cs.CY cs.AI cs.MA 81%

Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem

相对委托人、多元主义对齐与结构价值对齐问题

Travis LaCroix

机构 * Durham University(杜伦大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出价值对齐问题应被视为治理结构问题,通过三个轴框架分析对齐的成因,强调对齐是治理而非工程问题,需通过持续的制度过程管理。

Comments Accepted in the Ninth Annual ACM Conference on Fairness, Accountability, and Transparency (ACM FAccT) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17413 2026-04-21 cs.CY cs.AI 81%

The Open-Weight Paradox: Why Restricting Access to AI Models May Undermine the Safety It Seeks to Protect

开放权重悖论:为何限制访问AI模型可能反而削弱其安全防护

Vinicius Santana Gomes

机构 * Government of the Federal District of Brazil(巴西联邦区政府) School of Government of the Federal District of Brazil(巴西联邦区政府学校)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

AI总结 本文挑战开放权重AI模型治理的二元选择框架,提出无监管替代方案的访问限制可能转移而非减少风险,主张通过硬件层治理和多边机构架构提升AI安全。

Comments 23 pages, 2 figures, 1 table. Preprint also deposited at Zenodo (DOI: 10.5281/zenodo.19484877) on 2026-04-09. Licensed under CC BY 4.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11322 2026-04-14 cs.CL cs.AI 81%

Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations

LLMs是否了解工具无关性?解密工具调用中的结构对齐偏差

Yilong Liu, Xixun Lin, Pengfei Cao, Ge Zhang, Fang Fang, Yanan Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Computer Science and Technology, Donghua University(东华大学计算机科学与技术学院)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本文研究LLMs在面对无关工具时的调用偏差,提出SABEval数据集和Contrastive Attention Attribution方法,揭示结构对齐偏差的成因并提出缓解策略。

Comments Accepted to ACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13854 2026-03-27 cs.CY cs.AI 81%

Understanding the Process of Human-AI Value Alignment

理解人机价值对齐的过程

Jack McKinlay, Marina De Vos, Janina A. Hoffmann, Andreas Theodorou

机构 * University of Bath(巴斯大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 本文通过系统文献综述,探讨人工智能中价值对齐的核心方法与挑战,提出更精确的定义,识别六个关键主题以指导未来研究。

Comments 39 pages, 7 figures

Journal ref JAIR, Vol 85, Article 29 (March 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01214 2026-03-13 cs.CL cs.LG 81%

Reasoning Boosts Opinion Alignment in LLMs

推理提升大语言模型中的观点一致性

Frédéric Berdoz, Yann Billeter, Yann Vonlanthen, Roger Wattenhofer

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.LG

AI总结 本研究通过结构化推理提升大语言模型的观点一致性,验证了推理在改善观点建模中的有效性,但指出仍需进一步机制以减少偏见。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02420 2026-03-10 cs.CY cs.AI 81%

Slurry-as-a-Service: A Modest Proposal on Scalable Pluralistic Alignment for Nutrient Optimization

泥浆即服务:一种可扩展的多元对齐方案用于营养优化

Rachel Hong, Yael Eiger, Jevan Hutson, Os Keyes, William Agnew

机构 * ValueMulch

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出ValueMulch框架,通过多元对齐方法提升覆盖模型对社区规范的符合度,探讨伦理与技术对齐的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00078 2026-03-03 cs.CY cs.AI 81%

Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction

对齐不足以:人类-人工智能互动中的道德地位关系框架

Faezeh B. Pasandi, Hannah B. Pasandi

专题命中 AI治理与伦理 :alignment(title);trustworthy(abstract);分类 cs.AI、cs.CY

AI总结 本文提出Relate框架,通过关系能力和具身互动重新定义人工智能的道德资格,强调现有伦理词汇无法应对人类-人工智能互动的复杂性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22557 2026-02-27 cs.AI cs.LG 81%

CourtGuard: A Model-Agnostic Framework for Zero-Shot Policy Adaptation in LLM Safety

CourtGuard:一种用于大语言模型安全的模型无关框架,用于零样本策略适应

Umid Suleymanov, Rufiz Bayramov, Suad Gafarli, Seljan Musayeva, Taghi Mammadov, Aynur Akhundlu, Murat Kantarcioglu

机构 * Department of Computer Science, Virginia Tech(弗吉尼亚理工大学计算机科学系) School of IT(信息技术学院) Engineering, ADA University(工程学院,ADA大学) School of Law, ADA University(法学院,ADA大学)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 CourtGuard通过证据辩论框架实现大语言模型的安全零样本适应,超越传统方法在多个安全基准测试中的表现,并具备跨领域泛化和自动数据审计能力。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16438 2026-02-19 cs.LG cs.AI 81%

Intra-Fairness Dynamics: The Bias Spillover Effect in Targeted LLM Alignment

内在公平动态:定向大语言模型对齐中的偏见溢出效应

Eva Paraschou, Line Harder Clemmensen, Sneha Das

机构 * Department of Applied Mathematics and Computer Science, Technical University of Denmark(应用数学与计算机科学系,丹麦技术大学) Department of Mathematical Sciences, University of Copenhagen(数学科学系,哥本哈根大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了定向性别对齐对多敏感属性公平性的影响,发现改善某一属性公平性可能加剧其他属性的不平等,强调了多属性、情境感知公平性评估的重要性。

Comments Submitted to the BiAlign CHI Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22486 2026-02-02 cs.CY cs.AI 81%

AI Literacy, Safety Awareness, and STEM Career Aspirations of Australian Secondary Students: Evaluating the Impact of Workshop Interventions

澳大利亚中学生AI素养、安全意识与STEM职业抱负:评估工作坊干预的影响

Christian Bergh, Alexandra Vassar, Natasha Banks, Jessica Xu, Jake Renzella

机构 * University of New South Wales, Sydney(新南威尔士大学悉尼分校)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

AI总结 本研究评估了工作坊干预对澳大利亚中学生AI素养、安全意识和STEM职业兴趣的影响,发现干预提升了学生对AI的认知和兴趣,但需持续干预以影响长期职业抱负。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14298 2026-01-22 cs.CR cs.AI cs.CY 81%

Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)

大型语言模型(LLM)的信任、安全与伦理发展和部署的防护措施

Anjanava Biswas, Wrick Talukdar

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出了一种灵活自适应序列机制,结合信任和安全模块,用于实现大型语言模型开发和部署的安全防护。

Journal ref Journal Of Science & Technology, Vol. 4 No. 6 (2023), 55-82

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13565 2026-01-14 cs.CY cs.AI cs.HC 81%

Aligning Trustworthy AI with Democracy: A Dual Taxonomy of Opportunities and Risks

将可信AI与民主对齐:机会与风险的双重分类

Oier Mentxaka, Natalia Díaz-Rodríguez, Mark Coeckelbergh, Marcos López de Prado, Emilia Gómez, David Fernández Llorca, Enrique Herrera-Viedma, Francisco Herrera

机构 * Dept. of Computer Science and Artificial Intelligence, DaSCI, University of Granada, Spain(计算机科学与人工智能系,DaSCI,格拉纳达大学) Dept. of Philosophy, University of Vienna, Vienna, Austria(哲学系,维也纳大学) School of Engineering, Cornell University, Ithaca, NY, United States(工程学院,康奈尔大学) Dept. of Mathematics, Khalifa University of Science and Technology, Abu Dhabi, UAE(数学系,科学与技术大学,阿布扎赫德,阿联酋) ADIA Lab, Al Maryah Island, Abu Dhabi, UAE(ADIA实验室,阿布扎赫德,阿联酋) Joint Research Centre, European Commission, Seville, Spain(联合研究中心,欧洲委员会,塞维利亚,西班牙) Computer Engineering Dept., University of Alcalá, Alcalá de Henares, Spain(计算机工程系,阿尔卡拉大学)

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出双重分类框架,评估AI对民主的风险与机遇,为可信民主AI的发展提供规范性指导。

Comments 26 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11950 2025-12-16 cs.CY cs.AI 81%

Privacy Ethics Alignment in AI: A Stakeholder-Centric Framework for Ethical AI

AI隐私伦理对齐:面向伦理AI的利害相关者中心框架

Ankur Barthwal, Molly Campbell, Ajay Kumar Shrestha

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 本研究提出PEA-AI模型,通过分析不同利害相关者群体的隐私需求,构建以透明度和数字素养为核心的AI隐私治理框架。

Comments Peer reviewed publication at 10.3390/systems13060455

Journal ref Systems 2025, 13, 455

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06249 2025-12-11 cs.CL cs.AI 81%

TRepLiNa: Layer-wise CKA+REPINA Alignment Improves Low-Resource Machine Translation in Aya-23 8B

TRepLiNa:分层CKA+REPINA对齐改进Aya-23 8B低资源机器翻译

Toshiki Nakai, Ravi Kiran Chikkala, Lena Sophie Oberkircher, Nicholas Jennings, Natalia Skachkova, Tatiana Anikina, Jesujoba Oluwadara Alabi

机构 * Saarland University(萨尔兰大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 TRepLiNa通过结合CKA和REPINA实现分层对齐,提升低资源语言Aya-23 8B的机器翻译质量,尤其在数据稀缺情况下效果显著。

Comments It is work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00313 2025-11-25 cs.GT cs.AI cs.CL cs.MA 81%

Distributive Fairness in Large Language Models: Evaluating Alignment with Human Values

大语言模型中的分配公平性:评估与人类价值观的对齐

Hadi Hosseini, Samarth Khanna

机构 * Penn State University(宾夕法尼亚州立大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本文研究了大语言模型在分配公平性方面的表现,发现其与人类价值观存在偏差,并探讨了提升对齐性的策略。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17256 2025-11-24 cs.CY cs.CL 81%

Cross-cultural value alignment frameworks for responsible AI governance: Evidence from China-West comparative analysis

跨文化价值观对齐框架用于负责任的AI治理:来自中西方比较分析的证据

Haijiang Liu, Jinguang Gu, Xun Wu, Daniel Hershcovich, Qiaoling Xiao

机构 * Wuhan University of Science and Technology(武汉理工大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Copenhagen(哥本哈根大学) WUST-Madrid Complutense Institute(武汉理工大学-马德里康普顿斯研究所)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.CY

AI总结 本研究提出多层审计平台,评估中西方LLM跨文化价值观对齐,发现中国模型在多语言优化上更优,而西方模型存在美国偏见,Mistral系列在跨文化对齐上表现突出。

Comments Presented on Academic Conference "Technology for Good: Driving Social Impact" (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08592 2025-11-21 cs.LG cs.AI cs.ET 81%

Interpretability as Alignment: Making Internal Understanding a Design Principle

可解释性作为对齐:使内部理解成为设计原则

Aadit Sengupta, Pratinav Seth, Vinay Kumar Sankarapu

机构 * Lexsi Labs(Lexsi实验室)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出将可解释性作为设计原则,通过机理可解释性构建私营AI治理的基础设施,实现技术可靠性与制度问责的结合。

Comments Accepted at the first EurIPS Workshop on Private AI Governance

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09663 2025-11-14 cs.CY cs.AI cs.HC 81%

Alignment Debt: The Hidden Work of Making AI Usable

Cumi Oyemike, Elizabeth Akpan, Pierre Hervé-Berdys

机构 * YUX Design(YUX设计)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments 19 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01885 2025-11-06 cs.AI cs.LG q-bio.NC 81%

Mirror-Neuron Patterns in AI Alignment

Robyn Wyrick

机构 * Department of Computer Science University of Bath, United Kingdom(计算机科学系 英国巴斯大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 51 pages, Masters thesis. 10 tables, 7 figures, project data & code here: https://github.com/robynwyrick/mirror-neuron-frog-and-toad

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09586 2025-11-05 cs.HC cs.AI cs.CL 81%

ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs

Hua Shen, Tiffany Knearem, Reshmi Ghosh, Yu-Ju Yang, Nicholas Clark, Tanushree Mitra, Yun Huang

机构 * NYU Shanghai(纽约大学上海校区) New York University(纽约大学) University of Washington(华盛顿大学) MBZUAI Microsoft(微软) UIUC(伊利诺伊大学香槟分校)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00379 2025-11-04 cs.AI cs.CL 81%

Diverse Human Value Alignment for Large Language Models via Ethical Reasoning

Jiahao Wang, Songkai Xue, Jinghui Li, Xiaozhen Wang

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by AIES 2025, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09131 2025-10-24 cs.LG cs.AI stat.ML 81%

Fair Clustering via Alignment

Kunwoong Kim, Jihu Lee, Sangchul Park, Yongdai Kim

机构 * Department of Statistics, Seoul National University, Republic of Korea(统计系,首尔国立大学,大韩民国) School of Law, Seoul National University, Republic of Korea(法学院,首尔国立大学,大韩民国)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

Journal ref ICML 2025 (Forty-Second International Conference on Machine Learning)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09264 2025-09-30 cs.HC cs.AI cs.CL 81%

Position: Towards Bidirectional Human-AI Alignment

Hua Shen, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Ziqiao Ma, Savvas Petridis, Yi-Hao Peng, Li Qiwei, Sushrita Rakshit, Chenglei Si, Yutong Xie, Jeffrey P. Bigham, Frank Bentley, Joyce Chai, Zachary Lipton, Qiaozhu Mei, Rada Mihalcea, Michael Terry, Diyi Yang, Meredith Ringel Morris, Paul Resnick, David Jurgens

机构 * NYU Shanghai, New York University(纽约大学上海研究院,纽约大学) MBZUAI(穆桑大学人工智能研究所) Microsoft(微软公司) University of Michigan(密歇根大学) Apple(苹果公司) Google(谷歌公司) Carnegie Mellon University(卡内基梅隆大学) Stanford University(斯坦福大学) Google DeepMind(谷歌DeepMind)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025 Position Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11620 2025-09-16 cs.CL cs.CY 81%

AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment

Kun Li, Lai-Man Po, Hongzheng Yang, Xuyuan Xu, Kangcheng Liu, Yuzhi Zhao

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08858 2025-09-15 cs.CY cs.LG 81%

Decentralising LLM Alignment: A Case for Context, Pluralism, and Participation

Oriane Peter, Kate Devlin

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CY、cs.LG

Comments Accepted at AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09016 2025-09-11 cs.CL cs.LG 81%

A Survey on Training-free Alignment of Large Language Models

Birong Pan, Yongqi Li, Weiyu Zhang, Wenpeng Lu, Mayi Xu, Shen Zhou, Yuanyuan Zhu, Ming Zhong, Tieyun Qian

机构 * School of Computer Science, Wuhan University, Wuhan, China(武汉大学计算机学院) Faculty of Computer Science and Technology, Qilu University of Technology, Shandong, China(青岛科技大学计算机科学与技术学院) Zhongguancun Academy, Beijing, China(中关村学院)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments Accepted to EMNLP 2025 (findings), camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07022 2025-09-10 cs.CY cs.AI 81%

Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants

Pavan Reddy, Nithin Reddy

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

Comments 7 pages content, 1 page reference, 1 figure, Accepted at AAAI Fall Symposium Series

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05432 2025-08-08 cs.AI cs.CY 81%

Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI

Krzysztof Janowicz, Zilong Liu, Gengchen Mai, Zhangyu Wang, Ivan Majic, Alexandra Fortacz, Grant McKenzie, Song Gao

机构 * University of Vienna(维也纳大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Maine(缅因大学) McGill University(麦吉尔大学) University of Wisconsin(威斯康星大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏