arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-05-27 至 2026-05-27 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 4 篇

2512.11878 2026-05-27 cs.CY cs.CR 74%

A Technical Policy Blueprint for Trustworthy Decentralized AI

可信去中心化人工智能的技术政策蓝图

Hasan Kassem, Orion Banks, Omar Benjelloun, Sergen Cansiz, Brandon Edwards, Patrick Foley, Inken Hagestedt, Taeho Jung, Peter Kairouz, Marco Lorenzi, Peter Mattson, Prakash Moorthy, Ann K Novakowski, Michael O'Connor, Bruno Rodrigues, Holger Roth, Micah Sheller, Dimitris Stripelis, Renato Umeton, Marc Vesin, Wenbin Zhang, Mic Bowman, Alexandros Karargyris

专题命中 AI治理与伦理 :trustworthy(title);分类 cs.CY

AI总结 提出一种技术政策蓝图,通过将治理需求编码为策略即代码对象,并解耦策略验证与执行,实现去中心化AI系统的透明、可扩展和可验证治理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06708 2026-05-27 cs.LG cs.AI 62%

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights

通过模仿模型权重评估样本效用以实现高效数据选择

Tzu-Heng Huang, Manjot Bilkhu, John Cooper, Frederic Sala, Javier Movellan

机构 * Apple(苹果公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 提出基于梯度和几何的Mimic Score指标,通过Grad-Mimic框架在线重加权样本加速训练、离线构建数据过滤器,在六个图像数据集上提升数据效率和CLIP模型性能。

Comments This work appears in the Proceedings of the 43rd International Conference on Machine Learning (ICML 2026) and was selected as an Oral paper at the ICML 2025 DataWorld Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22648 2026-05-27 cs.AI cs.LG 62%

UCPO: Uncertainty-Aware Policy Optimization

UCPO:不确定性感知策略优化

Xianzhou Zeng, Jing Huang, Chunmei Xie, Gongrui Nan, Siye Chen, Mengyu Lu, Weiqi Xiong, Qixuan Zhou, Junhao Zhang, Qiang Zhu, Yadong Li, Xingzhong Xu

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 针对现有强化学习范式在不确定性奖励下存在的优势偏差和过度自信问题,提出三元优势解耦和动态不确定性奖励调整机制,显著提升模型在知识边界外的可靠性。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26870 2026-05-27 cs.MA cs.AI cs.HC 57%

Persistent AI Agents in Academic Research: A Single-Investigator Implementation Case Study

学术研究中的持久性AI智能体:单研究者实施案例研究

Anas H. Alzahrani

机构 * Department of Preventive Medicine and Public Health, Faculty of Medicine, King Abdulaziz University(预防医学与公共卫生系,医学院,国王阿卜杜勒阿齐兹大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 通过单研究者案例研究,分析了持久性AI智能体在真实学术环境中的架构、使用、产出和治理,发现缓存主导的工作流可能将经济单位从每token成本转向每完成工件成本。

Comments 19 pages, 2 figures, 3 main tables; supplementary appendix with 6 tables, 2 figures, and a reproducibility methods section. Describes 17 configured agents in a persistent research environment and introduces the PARE-M (Persistent Agentic Research Environment Measurement) framework

详情

展开后加载摘要…

URL PDF HTML 收藏