arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-04-09 至 2026-04-09 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 4 篇

2604.06215 2026-04-09 cs.CY cs.AI 73%

Governing frontier general-purpose AI in the public sector: adaptive risk management and policy capacity under uncertainty through 2030

在公共部门治理前沿通用人工智能:通过2030年的不确定性实现适应性风险管理与政策能力

Fabio Correa Xavier

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

AI总结 本文探讨公共部门在2030年前应对前沿通用人工智能的治理挑战,提出基于适应性风险管理、情景感知监管和社会技术转型的治理框架,强调政策能力提升与责任分配的重要性。

Comments 7 PAGES, 1 FIGURE

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06219 2026-04-09 cs.CY cs.AI 62%

From experimentation to engagement: on the paradox of participatory AI and power in contexts of forced displacement and humanitarian crises

从实验到参与:关于参与式AI与权力在流离失所和人道主义危机中的悖论

Stella Suge, Sarah W. Spencer, Nyalleng Moorosi, Helen McElhinney, Geoff Loane, Sue Black

机构 * FilmAid Kenya(肯尼亚电影援助组织) The Distributed AI Research Institute (DAIR)(分布式人工智能研究所) The CDAC Network(CDAC网络) Durham University(杜伦大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文探讨参与式AI在流离失所和人道主义危机中的局限性,指出其可能加剧'参与洗白'和算法伤害,强调需更严谨的方法和独立治理架构。

Comments This paper was submitted to the ACM FAccT conference in 2025 and is published here as a preprint in March 2026. The research was conducted in December 2024. Since submission, AI deployment across the humanitarian sector has accelerated without commensurate development of independent accountability

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06203 2026-04-09 cs.CY cs.AI 62%

Front-End Ethics for Sensor-Fused Health Conversational Agents: An Ethical Design Space for Biometrics

传感器融合健康对话代理的前端伦理:生物特征的伦理设计空间

Hansoo Lee, Rafael A. Calvo

机构 * Imperial College London(伦敦帝国理工学院) Korea Institute of Science and Technology(韩国科学技术研究院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文探讨了生物特征翻译的伦理设计,提出五个维度分析前端伦理风险,提出适应性披露作为安全机制,确保健康代理支持用户自主性。

Comments Accepted at the Proceedings of the CHI 2026 Workshop: Ethics at the Front-End

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09655 2026-04-09 cs.CR 50%

How Secure is Code Generated by ChatGPT?

ChatGPT生成的代码安全性如何?

Raphaël Khoury, Anderson R. Avila, Jacob Brunelle, Baba Mamadou Camara

专题命中 AI治理与伦理 :safety(abstract)

AI总结 研究评估了ChatGPT生成代码的安全性,探讨了通过提示提高安全性的方法及AI生成代码的伦理问题。

Journal ref 2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC) October 1-4, 2023, Oahu, Hawaii, USA

详情

展开后加载摘要…

URL PDF HTML 收藏