arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2603.24580 2026-03-26 cs.CL cs.AI cs.CY cs.IR cs.LG 70%

Retrieval Improvements Do Not Guarantee Better Answers: A Study of RAG for AI Policy QA

检索改进并不保证更好的答案:RAG在人工智能政策问答中的研究

Saahil Mathur, Ryan David Rittner, Vedant Ajit Thakur, Daniel Stuart Schiff, Tunazzina Islam

机构 * Department of Computer Science, Purdue University(计算机科学系,普渡大学) Department of Political Science, Purdue University(政治学系,普渡大学)

专题命中 AI治理与伦理 :DPO(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文研究了RAG在人工智能治理中的应用,发现领域特定微调虽提升检索指标,但未必改善问答性能,尤其在相关文档缺失时可能导致更自信的幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15900 2026-03-18 cs.NI cs.AI 70%

The Internet of Physical AI Agents: Interoperability, Longevity, and the Cost of Getting It Wrong

物理AI代理的互联网:互操作性、持久性与搞错的成本

Roberto Morabito, Mallik Tatipamula

机构 * EURECOM Ericsson Silicon Valley(爱立信硅谷)

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI

AI总结 本文探讨物理AI代理的互联网发展,提出设计原则以构建可靠、可进化和可信的代理系统,强调互操作性、信任和进化作为首要需求,避免技术与经济成本。

Comments A related version of this work is currently under review for publication in an IEEE magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05619 2026-03-04 cs.CR cs.LG 70%

LiteLMGuard: Seamless and Lightweight On-Device Prompt Filtering for Safeguarding Small Language Models against Quantization-induced Risks and Vulnerabilities

LiteLMGuard: 无缝且轻量级的设备端提示过滤,用于保护小型语言模型免受量化引发的风险和漏洞

Kalyan Nakka, Jimmy Dani, Ausmit Mondal, Nitesh Saxena

机构 * SPIES Research Lab, Dept. of CSE, Texas A&M University(SPIES研究实验室,计算机科学与工程系,德克萨斯大学阿马尔科分校)

专题命中 AI治理与伦理 :safety(abstract);jailbreak(abstract);分类 cs.LG

AI总结 LiteLMGuard通过设备端实时提示过滤,有效保护小型语言模型免受量化带来的安全风险。

Comments 18 pages, 19 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14606 2026-02-17 cs.MA cs.AI cs.CE 70%

Towards Selection as Power: Bounding Decision Authority in Autonomous Agents

迈向选择作为权力:自主代理中决策权威的边界

Jose Manuel de la Chica Rodriguez, Juan Manuel Vera Díaz

机构 * AI Lab, Grupo Santander Madrid, Spain(AI实验室,西班牙 Grupo Santander 马德里)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本文提出一种治理架构,通过机械机制限制自主代理的选择权,以防止确定性结果并提升可审计性,重新定义治理为受限制的因果权力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02582 2026-02-04 cs.AI cs.CL cs.CY cs.IR cs.LG cs.SE 70%

Uncertainty and Fairness Awareness in LLM-Based Recommendation Systems

大语言模型推荐系统中的不确定性与公平性意识

Chandan Kumar Sah, Xiaoli Lian, Li Zhang, Tony Xu, Syed Shazaib Shah

机构 * Beihang University(北京航空航天大学) McGill University(麦吉尔大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文研究了大语言模型推荐系统中不确定性与公平性的影响,提出了一种新的评估方法并揭示了个性化与公平性之间的权衡。

Comments Accepted at the Second Conference of the International Association for Safe and Ethical Artificial Intelligence, IASEAI26, 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01942 2026-02-03 cs.CR cs.AI 70%

Human Society-Inspired Approaches to Agentic AI Security: The 4C Framework

面向代理AI安全的人类社会启发方法:4C框架

Alsharif Abuadbba, Nazatul Sultan, Surya Nepal, Sanjay Jha

机构 * University of New South Wales, Sydney(新南威尔士大学悉尼分校)

专题命中 AI治理与伦理 :prompt injection(abstract);trustworthy(abstract);分类 cs.AI

AI总结 本文提出4C框架,通过人类社会治理理念,为代理AI安全提供多维度保护,强调行为完整性和意图,构建可信可控的AI系统。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00751 2026-02-03 cs.AI cs.SE 70%

Engineering AI Agents for Clinical Workflows: A Case Study in Architecture,MLOps, and Governance

为临床工作流程工程AI代理:架构、MLOps和治理的案例研究

Cláudio Lúcio do Val Lopes, João Marcus Pitta, Fabiano Belém, Gildson Alves, Flávio Vinícius Cruzeiro Martins

机构 * A3Data CEFET-MG

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI

AI总结 本文提出通过整合四个工程支柱构建可信临床AI系统,展示Maria平台在架构、MLOps和治理方面的创新实践。

Comments 9 pages, 5 figures 2026 IEEE/ACM 5th International Conference on AI Engineering - Software Engineering for AI}{April 12--13, 2026}{Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22621 2026-02-02 cs.CY 70%

Ethical Risks of Large Language Models in Medical Consultation: An Assessment Based on Reproductive Ethics

大语言模型在医疗咨询中的伦理风险:基于生殖伦理的评估

Hanhui Xu, Jiacheng Ji, Haoan Jin, Han Ying, Mengyue Wu

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

AI总结 本研究评估了大语言模型在生殖伦理咨询中的伦理风险,发现其在安全性和同理心方面存在严重缺陷,需加强推理和伦理依据能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13122 2026-01-21 cs.AI 70%

Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward

通用系统责任AI:概述、挑战与前行之路

Gourab K Patro, Himanshi Agrawal, Himanshu Gharat, Supriya Panigrahi, Nim Sherpa, Vishal Vaddina, Dagnachew Birru

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本文探讨了通用AI系统在责任AI方面的挑战,并提出C2V2需求以指导未来系统的开发。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03578 2026-01-08 cs.CL 70%

PsychEthicsBench: Evaluating Large Language Models Against Australian Mental Health Ethics

PsychEthicsBench: 评估大型语言模型在澳大利亚心理健康伦理中的表现

Yaling Shen, Stephanie Fong, Yiwen Jiang, Zimu Wang, Feilong Tang, Qingyang Xu, Xiangyu Zhao, Zhongxing Xu, Jiahe Liu, Jinpeng Hu, Dominic Dwyer, Zongyuan Ge

机构 * Monash University(墨尔本大学) University of Liverpool(利物浦大学) Hefei University of Technology(合肥工业大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 PsychEthicsBench通过多选和开放式任务评估LLMs在心理健康伦理方面的表现,揭示拒绝率作为伦理指标的不足,并指出领域微调可能损害伦理鲁棒性。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02749 2026-01-07 cs.AI 70%

The Path Ahead for Agentic AI: Challenges and Opportunities

代理AI的未来之路:挑战与机遇

Nadia Sibai, Yara Ahmed, Serry Sibaee, Sawsan AlHalawani, Adel Ammar, Wadii Boulila

机构 * Robotics and Internet-of-Things (RIOTU) Lab, Prince Sultan University, Riyadh, Saudi Arabia(机器人与物联网(RIOTU)实验室,普森大学,利雅得,沙特阿拉伯)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本文探讨了代理AI从被动生成到自主行动的转变,分析了其技术挑战与核心方法,并提出了实现自主行为的关键研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04354 2025-12-05 cs.LG cs.HC 70%

SmartAlert: Implementing Machine Learning-Driven Clinical Decision Support for Inpatient Lab Utilization Reduction

SmartAlert:基于机器学习的临床决策支持系统用于住院患者实验室利用减少

April S. Liang, Fatemeh Amrollahi, Yixing Jiang, Conor K. Corbin, Grace Y. E. Kim, David Mui, Trevor Crowell, Aakash Acharya, Sreedevi Mony, Soumya Punnathanam, Jack McKeown, Margaret Smith, Steven Lin, Arnold Milstein, Kevin Schulman, Jason Hom, Michael A. Pfeffer, Tho D. Pham, David Svec, Weihan Chu, Lisa Shieh, Christopher Sharp, Stephen P. Ma, Jonathan H. Chen

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.LG

AI总结 SmartAlert通过机器学习驱动的临床决策支持系统,减少住院患者重复实验室检测,实现15%的检测减少率。

Comments 22 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19362 2025-12-02 cs.CV cs.AI cs.CL cs.CY cs.LG 70%

LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences

LOTUS: 一种用于从质量到社会偏见和用户偏好的详细图像描述的排行榜

Yusuke Hirota, Boyi Li, Ryo Hachiuma, Yueh-Hua Wu, Boris Ivanovic, Yuta Nakashima, Marco Pavone, Yejin Choi, Yu-Chiang Frank Wang, Chao-Han Huck Yang

机构 * NVIDIA Research(NVIDIA研究)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 LOTUS是一种用于评估详细图像描述质量、偏见和社会偏见的排行榜,通过定制标准满足不同用户偏好,揭示了模型在不同评估标准上的表现差异。

Comments Accepted to ACL 2025. Leaderboard: huggingface.co/spaces/nvidia/lotus-vlm-bias-leaderboard

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19334 2025-11-25 cs.CY 70%

Normative active inference: A numerical proof of principle for a computational and economic legal analytic approach to AI governance

规范性主动推断:一种计算和经济法律分析方法用于AI治理的数值原理证明

Axel Constant, Mahault Albarracin, Karl J. Friston

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

AI总结 本文提出通过设计监管实现合法且规范敏感的AI行为,利用主动推断框架模拟自动驾驶场景,展示法律规范对AI代理决策的影响。

Comments 19 pages, 6 figures, 1 box

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03858 2025-11-19 cs.AI cs.ET cs.MA 70%

MI9: An Integrated Runtime Governance Framework for Agentic AI

Charles L. Wang, Trisha Singhal, Ameya Kelkar, Jason Tuo

机构 * Barclays, Model Risk Management(巴克莱银行,模型风险管理部门) Columbia University(哥伦比亚大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20606 2025-11-19 cs.CL 70%

Model Editing as a Double-Edged Sword: Steering Agent Ethical Behavior Toward Beneficence or Harm

Baixiang Huang, Zhen Tan, Haoran Wang, Zijie Liu, Dawei Li, Ali Payani, Huan Liu, Tianlong Chen, Kai Shu

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

Comments AAAI 2026 Oral. 14 pages (including appendix), 11 figures. Code, data, results, and additional resources are available at: https://model-editing.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10704 2025-11-19 cs.AI 70%

The Second Law of Intelligence: Controlling Ethical Entropy in Autonomous Systems

Samih Fadli

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 12 pages, 4 figures, 1 table, includes Supplementary Materials, simulation code on GitHub (https://github.com/AerisSpace/SecondLawIntelligence )

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16851 2025-11-17 cs.CR cs.CL 70%

Interpretable LLM Guardrails via Sparse Representation Steering

Zeqing He, Zhibo Wang, Huiyu Xu, Hejun Lin, Wenhui Zhang, Zhixuan Chu

机构 * The State Key Laboratory of Blockchain and Data Security, Zhejiang University, China(区块链与数据安全国家重点实验室,浙江大学,中国) School of Cyber Science and Technology, Zhejiang University, China(网络安全与技术学院,浙江大学,中国) College of Computer and Information Sciences, Fujian Agriculture and Forestry University, China(计算机与信息科学学院,福建农林大学,中国)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09606 2025-11-04 cs.CY 70%

Local US officials' views on the impacts and governance of AI: Evidence from 2022 and 2023 survey waves

Sophia Hatz, Noemi Dreksler, Kevin Wei, Baobao Zhang

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Journal ref PLoS One 20(10): e0332919, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00024 2025-11-04 cs.CY cs.AI cs.CL cs.LG stat.AP 70%

Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model

Haotian Hang, Yueyang Shen, Vicky Zhu, Jose Cruz, Michelle Li

机构 * University of Southern California(南加州大学) University of Michigan(密歇根大学) Babson College(巴布森学院) University of Connecticut(康涅狄格大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19327 2025-10-23 cs.MA cs.AI 70%

SORA-ATMAS: Adaptive Trust Management and Multi-LLM Aligned Governance for Future Smart Cities

Usama Antuley, Shahbaz Siddiqui, Sufian Hameed, Waqas Arif, Subhan Shah, Syed Attique Shah

机构 * organization= Department of Computer Science, National University of Computer \& Emerging Sciences , addressline= St-4 Sector 17-D On National Highway , city= Karachi , postcode= 75160 , state= , country= Pakistan organization= Balochistan University of Information Technology, Engineering organization= Department of Computer Science, Birmingham City University , addressline= STEAMhouse, Belmont Row , city= Birmingham , postcode= B4 7RQ , country= United Kingdom

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14053 2025-10-17 cs.AI 70%

Position: Require Frontier AI Labs To Release Small "Analog" Models

Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer, Philip Quirke

机构 * Frontier AI Labs(前沿人工智能实验室)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10998 2025-10-14 cs.CL cs.AI cs.CY cs.HC cs.LG 70%

ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios

Mahika Phutane, Hayoung Jung, Matthew Kim, Tanushree Mitra, Aditya Vashistha

机构 * Cornell University(康奈尔大学) Princeton University(普林斯顿大学) University of Washington(华盛顿大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 28 pages, 11 figures, 16 tables. In submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09871 2025-10-14 cs.CL 70%

CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs

Nafiseh Nikeghbal, Amir Hossein Kargaran, Jana Diesner

机构 * Technical University of Munich(慕尼黑技术大学) LMU Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06303 2025-10-07 cs.CY cs.AI cs.CL cs.LG 70%

On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions

Dang Nguyen, Chenhao Tan

机构 * Department of Computer Science University of Chicago(计算机科学系芝加哥大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 21 pages, 15 figures, 14 tables. Accepted as a conference paper at COLM 2025. Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22709 2025-09-30 cs.CY cs.SY eess.SY 70%

Trust and Transparency in AI: Industry Voices on Data, Ethics, and Compliance

Louise McCormack, Diletta Huyskes, Dave Lewis, Malika Bendechache

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21075 2025-09-26 cs.CY cs.AI cs.CL cs.DC cs.HC cs.LG 70%

Communication Bias in Large Language Models: A Regulatory Perspective

Adrian Kuenzler, Stefan Schmid

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13960 2025-08-20 cs.GT cs.AI 70%

A Mechanism for Mutual Fairness in Cooperative Games with Replicable Resources -- Extended Version

Björn Filter, Ralf Möller, Özgür Lütfü Özçep

机构 * Institute for Humanities-Centered AI (CHAI), University of Hamburg, Germany(人文中心人工智能研究所(CHAI),汉堡大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

Comments This paper is the extended version of a paper accepted at the European Conference on Artificial Intelligence 2025 (ECAI 2025), providing the proof of the main theorem in the appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07623 2025-08-13 cs.CL 70%

Optimizing Class-Level Probability Reweighting Coefficients for Equitable Prompting Accuracy

Ruixi Lin, Yang You

机构 * Department of Computer Science(计算机科学系) National University of Singapore(新加坡国立大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20014 2025-07-31 cs.CR cs.AI 70%

Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation

Joydeep Chandra, Satyam Kumar Navneet

机构 * Department of CST Tsinghua University Beijing, China(计算机科学与技术系 清华大学 北京中国) Department of CSE Chandigarh University Mohali, India(计算机科学与工程系 印度昌迪加尔大学 摩哈利)

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏