arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-04-23 至 2026-04-23 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 6 篇

2604.19752 2026-04-23 cs.MA cs.AI cs.CY 81%

Soft-Label Governance for Distributional Safety in Multi-Agent Systems

多智能体系统中的分布安全软标签治理

Aizierjiang Aiersilan, Raeli Savitt

机构 * The George Washington University(乔治·华盛顿大学) SWARM AI Safety(SWARM AI安全)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出SWARM框架,通过连续概率标签提升多智能体系统安全性和治理效果,展示软指标在检测代理游戏中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20805 2026-04-23 cs.CY cs.AI cs.MA 81%

Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem

相对委托人、多元主义对齐与结构价值对齐问题

Travis LaCroix

机构 * Durham University(杜伦大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出价值对齐问题应被视为治理结构问题,通过三个轴框架分析对齐的成因,强调对齐是治理而非工程问题,需通过持续的制度过程管理。

Comments Accepted in the Ninth Annual ACM Conference on Fairness, Accountability, and Transparency (ACM FAccT) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09871 2026-04-23 cs.AI cs.HC cs.LG 62%

Epistemology gives a Future to Complementarity in Human-AI Interactions

认识论为人类-人工智能互动中的互补性赋予未来

Andrea Ferrario, Alessandro Facchini, Juan M. Durán

机构 * SUPSI, Dalle Molle Institute for Artificial Intelligence (IDSIA)(SUPSI,达姆施塔特人工智能研究所(IDSIA)) Management in Networked and Digital Societies (MINDS) Department, Kozminski University(网络化与数字化社会管理系(MINDS)部,科米斯基大学) TU Delft(代尔夫特理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文通过认识论重构互补性,将其作为可靠知识过程的证据,提升人类-人工智能团队预测的可靠性,并提出设计与治理建议。

Comments Submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20199 2026-04-23 cs.CL 57%

All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG

所有语言都重要:理解并缓解多语言RAG中的语言偏差

Dan Wang, Guozhao Mo, Yafei Shi, Cheng Zhang, Bo Zheng, Boxi Cao, Xuanang Chen, Yaojie Lu, Hongyu Lin, Ben He, Xianpei Han, Le Sun

机构 * Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所信息处理实验室) University of Chinese Academy of Sciences(中国科学院大学) MYbank, AntGroup(蚂蚁集团MYbank)

专题命中 AI治理与伦理 :alignment(abstract_cn);分类 cs.CL

AI总结 本文研究多语言RAG中的语言偏差问题,提出LAURA方法,通过多语言证据排名与生成效用对齐,有效缓解语言偏差并提升性能。

Comments ACL 2026 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21293 2026-04-23 cs.AI cs.HC 57%

Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles

理解人工智能可信性:AIES与FAccT文章的综述

Siddharth Mehrotra, Jin Huang, Xuelong Fu, Roel Dobbe, Clara I. Sánchez, Maarten de Rijke

机构 * University of Amsterdam \& Delft University of Technology The Netherlands University of Cambridge United Kingdom University of Amsterdam The Netherlands Delft University of Technology The Netherlands University of Amsterdam \& Delft University of Technology University of Cambridge University of Amsterdam Delft University of Technology

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 本文通过综述AIES和FAccT会议论文,探讨人工智能可信性的概念、测量与验证方法,揭示技术与社会维度的不足,提出跨学科研究的重要性。

Comments Submitted to Journal of Artificial Intelligence Research (JAIR)

Journal ref Journal of Artificial Intelligence Research (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03396 2026-04-23 cs.CL 57%

Breaking the Assistant Mold: Modeling Behavioral Variation in LLM Based Procedural Character Generation

打破助手模式:在基于大语言模型的程序化角色生成中建模行为变异

Maan Qraitem, Kate Saenko, Bryan A. Plummer

机构 * Boston University(波士顿大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文提出PersonaWeaver框架,通过解耦世界观构建与行为构建,生成更具多样性和戏剧张力的角色。

详情

展开后加载摘要…

URL PDF HTML 收藏