arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-05-11 至 2026-05-11 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 3 篇

2601.13372 2026-05-11 cs.CY 80%

Semantic Alignment Between Normative Theories of Ethics and the European Union Artificial Intelligence Act: A Transformer-Based Semantic Textual Similarity Analysis

伦理规范理论与欧盟人工智能法案之间的语义对齐:基于Transformer的语义文本相似性分析

Mehmet Murat Albayrakoglu, Mehmet Nafiz Aydin

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CY

AI总结 本文通过Transformer模型分析伦理理论与欧盟AI法案的语义对齐,发现义务伦理学在法案两部分中具有最高相似性。

Comments 18 pages, 5 tables, 3 figures; the concept of alignment introduced as an indication of influence

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00398 2026-05-11 cs.CY cs.AI 62%

A Study on the Framework for Evaluating the Ethics and Trustworthiness of Generative AI

对生成式AI伦理性和可信度评估框架的研究

Cheonsu Jeong, Seunghyun Lee, Seonhee Jeong, Sungsu Kim

机构 * Hyper Automation Team, SAMSUNG SDS(三星SDS超自动化团队) Digital CRM Team, SAMSUNG SDS(三星SDS数字客户关系管理团队)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文研究生成式AI的伦理和可信度评估框架,提出系统评估方法,涵盖公平性、透明度等关键维度,并分析不同国家的AI伦理政策。

Comments 22 pages, 3 figures, 6 tables

Journal ref Artificial Intelligence and Applications, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.04321 2026-05-11 cs.CY cs.HC 57%

AI and Suicide Prevention: A Cross-Sector Primer

人工智能与自杀预防:跨行业的入门指南

Emily Saltz, Claire R. Leibowicz

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本文探讨了人工智能聊天机器人在心理健康支持中的应用,分析了其在临床验证、标准制定和监管协调方面的不足,并提出跨行业协作的必要性。

Comments 38 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏