REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.CL、cs.LG
Comments refactor
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.CL、cs.LG
Comments refactor
机构 * Dataplicada
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI
机构 * Brussels Centre for Language Studies, Vrije Universiteit Brussel(布鲁塞尔语言研究中心,布鲁塞尔自由大学) ; Université de Montréal & Mila - Quebec AI Institute(蒙特利尔大学及魁北克人工智能研究所)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
Comments Preprint under review at Computational Linguistics. Accepted with minor revisions (10/10/2025); second round
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
Comments Updated to the peer-reviewed version accepted and published in Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)
Journal ref Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
机构 * University of North Texas(北卡罗来纳大学达顿分校) ; Davidson College(戴维森学院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Istituto di Scienza e Tecnologie dell’Informazione, Consiglio Nazionale delle Ricerche(信息科学与技术研究所,国家研究理事会)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI
机构 * AWS Responsible AI(AWS负责任人工智能)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI
Comments 24 pages with 3 figures, to appear in Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM '25)
机构 * Department of Computer Science University of Technology Nuremberg(计算机科学系图腾技术大学纽伦堡)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 23 pages, 12 figures
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Technische Universität Dresden(德累斯顿技术大学) ; Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI)(可扩展数据与人工智能研究中心(ScaDS.AI))
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
机构 * University of Maryland(马里兰大学) ; Capital One
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.LG
Comments 22 Pages
机构 * Department of Information Systems, W.P. Carey School of Business, Arizona State University, Tempe, AZ, USA(亚利桑那州立大学信息系统系,W.P. Carey商学院,Tempe分校) ; Department of Computer Science, Cornell University, Ithaca, NY, USA(康奈尔大学计算机科学系) ; Graduate School of Management, University of California Davis, Davis, CA, USA(加州大学戴维斯分校管理研究生院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
Comments Accepted at PNAS Nexus
Journal ref PNAS Nexus 2025
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
机构 * Vrije University Amsterdam(荷兰阿姆斯特丹自由大学) ; Tri-institutional Center for Translational Research in Neuroimaging(转化神经影像研究联合中心) ; Emory University(埃默里大学) ; Key Laboratory of Genetic Evolution and Animal Models(遗传进化与动物模型重点实验室) ; Kunming Institute of Zoology(昆明动物研究所) ; Chinese Academy of Sciences Kunming(中国科学院昆明分院) ; Department of Psychiatry, Amsterdam UMC, University of Amsterdam(阿姆斯特丹大学精神病科) ; Department of Physics and Technology, UiT The Arctic University of Norway(北极大学挪威理工学院物理与技术系)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
Comments This manuscript has been accepted by Biomedical Signal Processing and Control and the code is available at https://github.com/TianzhengHU/BrainIB_coding/tree/main/BrainIB_GIB
机构 * Senior Research Fellow, University of Kent, UK(肯特大学高级研究员) ; Digital Child Safety Expert(数字儿童安全专家) ; Vice President of Data Science, Thorn(数据科学副总裁,Thorn)
专题命中 AI治理与伦理 :harmlessness(abstract);分类 cs.AI、cs.CY
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
机构 * Faculty of Computing - Federal University of Mato Grosso do Sul(计算机学院 - 短暂戈亚那联邦大学)
专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.LG
机构 * Department of Computer Science Virginia Tech(计算机科学系弗吉尼亚理工大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
Comments Accepted at 2025 ASEE Annual Conference & Exposition
机构 * The Ohio State University(俄亥俄州立大学)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments Add GPT 5 experiments
机构 * TU Delft(代尔夫特理工大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
Comments Proceeding of The British Academy of Management Conference 2025, University of Kent, UK
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
Comments 6 pages, no figures
Journal ref Nature, 644 (8075), 2025, 38-40
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY
Comments 10 pages, 5 figures. Accepted to the Workshop on Multimodal Continual Learning (MCL) at ICCV 2025. @2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), ICCV's 2025
机构 * University of Ioannina(伊奥安纳大学) ; Archimedes, Athena Research Center(阿基米德研究所) ; Boston University(波士顿大学)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments ECML PKDD 2025