arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-05-05 至 2026-05-05 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 3 篇

2605.01147 2026-05-05 cs.AI 89%

Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

位置:代理AI的安全性与公平性取决于交互拓扑,而非模型规模或对齐

Tanav Singh Bajaj, Nikhil Singh, Karan Anand, Eishkaran Singh

机构 * Department of Computer Science, University of British Columbia, Vancouver, Canada(英属哥伦比亚大学计算机科学系) Department of Artificial Intelligence, IIT Hyderabad, Hyderabad, India(印度海得拉巴理工学院人工智能系) Amazon, Delhi, India(印度德里亚马逊公司)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(title,abstract);AI safety(abstract);分类 cs.AI

AI总结 本文指出代理AI的安全性由交互拓扑决定,而非模型规模或对齐程度。研究揭示了顺序不稳定、信息级联和功能崩溃等拓扑驱动的问题,并强调需通过动态系统视角评估安全性和公平性。

Comments 18 pages, 8 figures. Position paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01229 2026-05-05 cs.LG cs.CL 62%

Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation

大规模多语言神经机器翻译中的注意力 sinks:发现、分析与缓解

Hillary Mutisya, John Mugane

机构 * Harvard University(哈佛大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 研究发现神经机器翻译中跨注意力模式存在注意力 sinks,非内容token占据大部分注意力质量,影响相似性评估,提出内容过滤方法以恢复语言信号。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01451 2026-05-05 cs.CL 61%

Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models

对基于AI的紧急警务调度中的种族偏见进行审计:对十一种大型语言模型的跨语言评估

William Guey, Wei Zhang, Pierrick Bougault, Yi Wang, Bertan Ucar, Vitor D. de Moura, José O. Gomes

机构 * Department of Industrial Engineering, Tsinghua University(清华大学工业工程系) School of Social Sciences, Tsinghua University(清华大学社会科学部) Department of Industrial Engineering, Federal University of Rio de Janeiro(里约热内卢联邦大学工业工程系)

专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.CL

AI总结 本文通过跨语言框架评估11种模型,在19800个输出中发现当事件严重性模糊时种族偏见系统性出现,但当操作优先级由通话内容确定时偏见消失。偏见程度因种族轴而异,宗教外观影响最大,性别次之,种族最小。语言间偏见转移不一致,性别偏见在中文中放大,种族偏见在英文中更明显。

Comments 26 pages, 7 figures. Submitted to Humanities and Social Sciences Communications (Nature) collection on Artificial Intelligence and Emerging Technologies in Public Safety. Code and data: https://github.com/williamguey/llmdispatchbias

详情

展开后加载摘要…

URL PDF HTML 收藏