arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2501.03262 2025-11-11 cs.CL cs.LG 62%

REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Jian Hu, Jason Klein Liu, Haotian Xu, Wei Shen

专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.CL、cs.LG

Comments refactor

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05766 2025-11-11 cs.AI cs.CL econ.GN q-fin.EC 62%

Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs

Felipe Valencia-Clavijo

机构 * Dataplicada

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03980 2025-11-07 cs.AI cs.CL 62%

LLMs and Cultural Values: the Impact of Prompt Language and Explicit Cultural Framing

Bram Bulté, Ayla Rigouts Terryn

机构 * Brussels Centre for Language Studies, Vrije Universiteit Brussel(布鲁塞尔语言研究中心,布鲁塞尔自由大学) Université de Montréal & Mila - Quebec AI Institute(蒙特利尔大学及魁北克人工智能研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Preprint under review at Computational Linguistics. Accepted with minor revisions (10/10/2025); second round

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17228 2025-11-06 cs.CY cs.AI 62%

Survey on AI Ethics: A Socio-technical Perspective

Dave Mbiazi, Meghana Bhange, Maryam Babaei, Ivaxi Sheth, Patrik Kenfack, Samira Ebrahimi Kahou

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments Updated to the peer-reviewed version accepted and published in Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)

Journal ref Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10603 2025-10-31 cs.CY cs.AI 62%

Toward a Public and Secure Generative AI: A Comparative Analysis of Open and Closed LLMs

Jorge Machado

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25218 2025-10-30 cs.CY cs.AI 62%

Human Resilience in the AI Era -- What Machines Can't Replace

Shaoshan Liu, Anina Schwarzenbach, Yiyu Shi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22823 2025-10-28 cs.CL cs.AI 62%

Cross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP

Poli Nemkova, Amrit Adhikari, Matthew Pearson, Vamsi Krishna Sadu, Mark V. Albert

机构 * University of North Texas(北卡罗来纳大学达顿分校) Davidson College(戴维森学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16836 2025-10-27 cs.CL cs.AI 62%

Misspellings in Natural Language Processing: A survey

Gianluca Sperduti, Alejandro Moreo

机构 * Istituto di Scienza e Tecnologie dell’Informazione, Consiglio Nazionale delle Ricerche(信息科学与技术研究所,国家研究理事会)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20782 2025-10-24 cs.CL cs.AI 62%

A Use-Case Specific Dataset for Measuring Dimensions of Responsible Performance in LLM-generated Text

Alicia Sagae, Chia-Jung Lee, Sandeep Avula, Brandon Dang, Vanessa Murdock

机构 * AWS Responsible AI(AWS负责任人工智能)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments 24 pages with 3 figures, to appear in Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19733 2025-10-24 cs.CL cs.LG 62%

Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning

M. H. I. Abdalla, Zhipin Wang, Christian Frey, Steffen Eger, Josif Grabocka

机构 * Department of Computer Science University of Technology Nuremberg(计算机科学系图腾技术大学纽伦堡)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18193 2025-10-23 cs.AI cs.CV cs.LG stat.ML 62%

FST.ai 2.0: An Explainable AI Ecosystem for Fair, Fast, and Inclusive Decision-Making in Olympic and Paralympic Taekwondo

Keivan Shariatmadar, Ahmad Osman, Ramin Ray, Kisam Kim

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 23 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14106 2025-10-17 cs.AI cs.CL cs.GT 62%

Generating Fair Consensus Statements with Social Choice on Token-Level MDPs

Carter Blair, Kate Larson

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08236 2025-10-17 cs.LG cs.AI 62%

The Hidden Bias: A Study on Explicit and Implicit Political Stereotypes in Large Language Models

Konrad Löhr, Shuzhou Yuan, Michael Färber

机构 * Technische Universität Dresden(德累斯顿技术大学) Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI)(可扩展数据与人工智能研究中心(ScaDS.AI))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09634 2025-10-14 cs.CY cs.AI 62%

Responsible AI Adoption in the Public Sector: A Data-Centric Taxonomy of AI Adoption Challenges

Anastasija Nikiforova, Martin Lnenicka, Ulf Melin, David Valle-Cruz, Asif Gill, Cesar Casiano Flores, Emyana Sirait, Mariusz Luterek, Richard Michael Dreyling, Barbora Tesarova

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07613 2025-10-10 cs.CL cs.AI 62%

Vocabulary embeddings organize linguistic structure early in language model training

Isabel Papadimitriou, Jacob Prince

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02563 2025-10-08 cs.LG cs.CL 62%

DynaGuard: A Dynamic Guardian Model With User-Defined Policies

Monte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah, Joseph Vincent, Chirag Jain, Melissa Kazemi Rad, C. Bayan Bruss, Ashwinee Panda, Tom Goldstein

机构 * University of Maryland(马里兰大学) Capital One

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.LG

Comments 22 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10659 2025-10-07 cs.SI cs.AI cs.CL cs.MA 62%

Network Formation and Dynamics Among Multi-LLMs

Marios Papachristou, Yuan Yuan

机构 * Department of Information Systems, W.P. Carey School of Business, Arizona State University, Tempe, AZ, USA(亚利桑那州立大学信息系统系,W.P. Carey商学院,Tempe分校) Department of Computer Science, Cornell University, Ithaca, NY, USA(康奈尔大学计算机科学系) Graduate School of Management, University of California Davis, Davis, CA, USA(加州大学戴维斯分校管理研究生院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at PNAS Nexus

Journal ref PNAS Nexus 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03368 2025-10-07 cs.CY cs.AI 62%

An Adaptive Responsible AI Governance Framework for Decentralized Organizations

Kiana Jafari Meimandi, Anka Reuel, Gabriela Aranguiz-Dias, Hatim Rahama, Ala-Eddine Ayadi, Xavier Boullier, Jérémy Verdo, Louis Montanie, Mykel Kochenderfer

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03004 2025-10-06 cs.LG cs.AI 62%

BrainIB++: Leveraging Graph Neural Networks and Information Bottleneck for Functional Brain Biomarkers in Schizophrenia

Tianzheng Hu, Qiang Li, Shu Liu, Vince D. Calhoun, Guido van Wingen, Shujian Yu

机构 * Vrije University Amsterdam(荷兰阿姆斯特丹自由大学) Tri-institutional Center for Translational Research in Neuroimaging(转化神经影像研究联合中心) Emory University(埃默里大学) Key Laboratory of Genetic Evolution and Animal Models(遗传进化与动物模型重点实验室) Kunming Institute of Zoology(昆明动物研究所) Chinese Academy of Sciences Kunming(中国科学院昆明分院) Department of Psychiatry, Amsterdam UMC, University of Amsterdam(阿姆斯特丹大学精神病科) Department of Physics and Technology, UiT The Arctic University of Norway(北极大学挪威理工学院物理与技术系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments This manuscript has been accepted by Biomedical Signal Processing and Control and the code is available at https://github.com/TianzhengHU/BrainIB_coding/tree/main/BrainIB_GIB

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02978 2025-10-06 cs.CY cs.AI cs.HC 62%

AI Generated Child Sexual Abuse Material -- What's the Harm?

Caoilte Ó Ciardha, John Buckley, Rebecca S. Portnoff

机构 * Senior Research Fellow, University of Kent, UK(肯特大学高级研究员) Digital Child Safety Expert(数字儿童安全专家) Vice President of Data Science, Thorn(数据科学副总裁,Thorn)

专题命中 AI治理与伦理 :harmlessness(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22774 2025-10-06 cs.AI cs.CY 62%

Bridging Ethical Principles and Algorithmic Methods: An Alternative Approach for Assessing Trustworthiness in AI Systems

Michael Papademas, Xenia Ziouvelou, Antonis Troumpoukis, Vangelis Karkaletsis

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26584 2025-10-01 cs.AI cs.IR cs.LG cs.SE 62%

Fairness Testing in Retrieval-Augmented Generation: How Small Perturbations Reveal Bias in Small Language Models

Matheus Vinicius da Silva de Oliveira, Jonathan de Andrade Silva, Awdren de Lima Fontao

机构 * Faculty of Computing - Federal University of Mato Grosso do Sul(计算机学院 - 短暂戈亚那联邦大学)

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07931 2025-10-01 cs.CY cs.AI 62%

Educating a Responsible AI Workforce: Piloting a Curricular Module on AI Policy in a Graduate Machine Learning Course

James Weichert, Hoda Eldardiry

机构 * Department of Computer Science Virginia Tech(计算机科学系弗吉尼亚理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Accepted at 2025 ASEE Annual Conference & Exposition

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16355 2025-09-29 cs.LG cs.AI 62%

How Strategic Agents Respond: Comparing Analytical Models with LLM-Generated Responses in Strategic Classification

Tian Xie, Pavan Rauch, Xueru Zhang

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Add GPT 5 experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17805 2025-09-29 cs.CY cs.AI 62%

Biospheric AI

Marcin Korecki

机构 * TU Delft(代尔夫特理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10653 2025-09-16 cs.CY cs.AI 62%

SCOR: A Framework for Responsible AI Innovation in Digital Ecosystems

Mohammad Saleh Torkestani, Taha Mansouri

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Proceeding of The British Academy of Management Conference 2025, University of Kent, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10289 2025-09-15 cs.CY cs.AI 62%

We Need a New Ethics for a World of AI Agents

Iason Gabriel, Geoff Keeling, Arianna Manzini, James Evans

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 6 pages, no figures

Journal ref Nature, 644 (8075), 2025, 38-40

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02773 2025-09-15 cs.CY cs.AI econ.GN q-fin.EC 62%

Web3 x AI Agents: Landscape, Integrations, and Foundational Challenges

Yiming Shen, Jiashuo Zhang, Zhenzhe Shao, Wenxuan Luo, Yanlin Wang, Ting Chen, Zibin Zheng, Jiachi Chen

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08829 2025-09-12 cs.CY cs.AI cs.IR 62%

PerFairX: Is There a Balance Between Fairness and Personality in Large Language Model Recommendations?

Chandan Kumar Sah

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 10 pages, 5 figures. Accepted to the Workshop on Multimodal Continual Learning (MCL) at ICCV 2025. @2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), ICCV's 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22591 2025-09-09 cs.LG cs.AI stat.ME 62%

FACEGroup: Feasible and Actionable Counterfactual Explanations for Group Fairness

Christos Fragkathoulas, Vasiliki Papanikou, Evaggelia Pitoura, Evimaria Terzi

机构 * University of Ioannina(伊奥安纳大学) Archimedes, Athena Research Center(阿基米德研究所) Boston University(波士顿大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments ECML PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏