arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1832 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1832 篇

2602.19682 2026-02-24 cs.CY 57%

Beyond the Binary: A nuanced path for open-weight advanced AI

超越二元:开放权重高级AI的细致路径

Bengüsu Özcan, Alex Petropoulos, Max Reddel

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

AI总结 本文提出了一种基于安全评估的分层模型发布方法,旨在解决开放权重高级AI模型在安全性和监管方面的挑战。

Comments This publication was originally designed and optimised for web and published on cfg.eu. Minor formatting differences may appear in this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07754 2026-02-24 cs.AI cs.HC 57%

Humanizing AI Grading: Student-Centered Insights on Fairness, Trust, Consistency and Transparency

让AI评分更人性化:以学生为中心的公平性、信任、一致性与透明性洞察

Bahare Riahi, Viktoriia Storozhevykh, Veronica Catete

机构 * North Carolina State University(北卡罗来纳州立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 本研究通过比较AI与人工评分反馈,探讨学生对AI评分系统在公平性、信任、一致性与透明性方面的看法,并提出人本化AI的设计原则。

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18535 2026-02-24 cs.SD cs.AI 57%

Fairness-Aware Partial-label Domain Adaptation for Voice Classification of Parkinson's and ALS

面向语音分类的公平性意识部分标签领域适应

Arianna Francesconi, Zhixiang Dai, Arthur Stefano Moscheni, Himesh Morgan Perera Kanattage, Donato Cappetta, Fabio Rebecchi, Paolo Soda, Valerio Guarrasi, Rosa Sicilia, Mary-Anne Hartley

机构 * organization= School of Computer Communication Sciences, EPFL (\'Ecole polytechnique f\'ed\'erale de Lausanne) , city= Lausanne , country= Switzerland organization= Eustema S.p.A., Research Development Centre , city= Naples , country= Italy organization= UniCamillus-Saint Camillus International University of Health Sciences , city= Rome , country= Italy organization= Department of Diagnostics Intervention, Radiation Physics, Biomedical Engineering, Umeå University , city= Umeå , country= Sweden

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种融合域泛化和对抗对齐的框架,用于在部分标签不匹配和公平性约束下实现帕金森病和肌萎缩侧索硬化症的统一语音分类。

Comments 7 pages, 1 figure. Submitted to Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18461 2026-02-24 cs.CY 57%

Toward Self-Driving Universities: Can Universities Drive Themselves with Agentic AI?

迈向自动驾驶大学:大学能否通过代理AI实现自我驱动?

Anis Koubaa

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本研究提出通过代理AI实现高等教育机构的自主性框架,旨在自动化行政、学术和质量保证流程,减少教师文书工作时间,提升教育质量和研究生产力。

Journal ref Springer Book: Next Generation AI-Driven Education - 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17919 2026-02-23 cs.CY cs.HC 57%

Visual Anthropomorphism Shifts Evaluations of Gendered AI Managers

视觉人化影响对性别化AI管理者评价

Ruiqing Han, Hao Cui, Taha Yasseri

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

AI总结 研究发现,文本描述中能力信息可缓解对AI管理者的负面评价,而视觉人化会引发性别偏见,表明表示方式影响性别刻板印象的激活。

Comments Preprint, Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16553 2026-02-19 cs.CY 57%

Agentic AI, Medical Morality, and the Transformation of the Patient-Physician Relationship

代理AI、医疗道德与患者-医生关系的变革

Robert Ranisch, Sabine Salloch

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

AI总结 本文探讨代理AI如何通过重塑患者-医生关系改变医疗道德,呼吁在广泛应用前融入伦理考量。

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12680 2026-02-16 stat.ML cs.LG 57%

A Regularization-Sharpness Tradeoff for Linear Interpolators

线性插值器的正则化-尖锐性权衡

Qingyi Hu, Liam Hodgkinson

机构 * School of Mathematics and Statistics(数学与统计学学院) University of Melbourne(墨尔本大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本文提出了一种针对过参数化线性回归的正则化-尖锐性权衡,通过ℓ^p惩罚分解选择惩罚为正则化项和几何尖锐性项,验证了其在现实数据中的有效性。

Comments 29 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11301 2026-02-13 cs.AI cs.CR 57%

The PBSAI Governance Ecosystem: A Multi-Agent AI Reference Architecture for Securing Enterprise AI Estates

PBSAI治理生态系统:一种多智能体AI参考架构,用于保障企业AI领地安全

John M. Willis

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 PBSAI提出了一种多智能体参考架构,用于保障企业AI领地的安全,通过十二个领域分类法和有限智能体家族实现责任划分,结合分析监控、协调防御等技术,确保系统安全与可追溯性。

Comments 43 pages, plus 12 pages of appendices. One Figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08193 2026-02-12 cs.AI 57%

Measuring What Matters: The AI Pluralism Index

衡量重要性:AI多元主义指数

Rashid Mushkani

机构 * Université de Montréal(蒙特利尔大学) Mila – Québec AI Institute(魁北克人工智能研究所)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 本文提出AI多元主义指数,用于衡量人工智能系统在治理、包容性和透明度方面的多元实践,旨在引导激励向多元主义方向发展。

Comments Proceedings of the International Association for Safe & Ethical AI (IASEAI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12365 2026-02-12 cs.CL cs.DB 57%

Advances in LLMs with Focus on Reasoning, Adaptability, Efficiency and Ethics

大语言模型的进展:聚焦推理、适应性、效率和伦理

Asifullah Khan, Muhammad Zaeem Khan, Aleesha Zainab, Saleha Jamshed, Sadia Ahmad, Kaynat Khatib, Faria Bibi, Abdul Rehman

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文综述了大语言模型在推理、适应性、效率和伦理方面的进展,探讨了关键技术和挑战,提出未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06218 2026-02-11 cs.CV cs.LG 57%

Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings

跨模态冗余与视觉-语言嵌入的几何学

Grégoire Dhimoïla, Thomas Fel, Victor Boutin, Agustin Picard

机构 * Brown University(布朗大学) ENS Paris Saclay(巴黎萨克雷大学) IRT Saint Exupéry(IRT圣埃克苏佩里) Kempner Institute, Harvard University(哈佛大学凯姆纳研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本文通过等能假设和对齐稀疏自编码器,揭示了视觉-语言模型中跨模态对齐的几何结构,发现稀疏双模态原子承载了跨模态对齐信号,单模态原子解释了模态差距,去除单模态原子可消除差距而不影响性能。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19186 2026-02-11 stat.ML cs.LG 57%

Double Fairness Policy Learning: Integrating Action Fairness and Outcome Fairness in Decision-making

双公平性政策学习:在决策中整合行动公平性与结果公平性

Zeyu Bian, Lan Wang, Chengchun Shi, Zhengling Qi

机构 * Department of Statistics(统计系) Florida State University(佛罗里达州立大学) Department of Management Science(管理科学系) University of Miami(迈阿密大学) London School of Economics and Political Science(伦敦政治经济学院) Department of Decision Sciences(决策科学系) George Washington University(乔治华盛顿大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

AI总结 本文提出双公平性学习框架,通过整合行动公平与结果公平,提升决策中的公平性并最小化价值损失。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06107 2026-02-09 cs.AI 57%

Jackpot: Optimal Budgeted Rejection Sampling for Extreme Actor-Policy Mismatch Reinforcement Learning

Jackpot: 为极端演员-策略不匹配强化学习的最优预算拒绝采样

Zhuoming Chen, Hongyi Liu, Yang Zhou, Haizhong Zheng, Beidi Chen

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 Jackpot通过最优预算拒绝采样方法,有效减少rollout模型与策略之间的分布差异,提升大语言模型强化学习的训练稳定性与效率。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03334 2026-02-04 cs.CY 57%

The Personality Trap: How LLMs Embed Bias When Generating Human-Like Personas

人格陷阱:大语言模型在生成类人人格时嵌入偏见的方式

Jacopo Amidei, Gregorio Ferreira, Mario Muñoz Serrano, Rubén Nieto, Andreas Kaltenbrunner

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本文研究了大语言模型在生成类人人格时嵌入WEIRD偏见的问题,揭示了LLMs在生成合成人口时可能带来的刻板印象和风险。

Comments 26 pages, 2 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02170 2026-02-03 cs.MA cs.AI 57%

Self-Evolving Coordination Protocol in Multi-Agent AI Systems: An Exploratory Systems Feasibility Study

多智能体AI系统中的自演化协调协议:一种探索性系统可行性研究

Jose Manuel de la Chica Rodriguez, Juan Manuel Vera Díaz

机构 * AI Lab, Grupo Santander Madrid, Spain(西班牙桑坦德集团AI实验室)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 本研究探讨了自演化协调协议在多智能体系统中的可行性,通过实验展示有限自我修改在满足形式约束下的技术实现可能性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00816 2026-02-03 stat.ML cs.LG 57%

Hessian Spectral Analysis at Foundation Model Scale

基础模型规模下的Hessian谱分析

Diego Granziol, Khurshid Juarev

机构 * Mathematical Institute, University of Oxford, UK(牛津大学数学研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本研究在大规模基础模型上实现了Hessian谱的准确分析,揭示了块对角曲率近似在中等规模LLM中的失效问题,展示了谱探测的计算效率与实际应用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00300 2026-02-03 cs.CL 57%

Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models

Faithful-Patchscopes: 理解和缓解大语言模型隐藏表示解释中的模型偏差

Xilin Gong, Shu Yang, Zehua Cao, Lynne Billard, Di Wang

机构 * University of Georgia(佐治亚大学) King Abdullah University of Science(国王阿卜杜勒-阿齐兹大学) Hong Kong Center for Construction Robotics(香港建筑机器人中心)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文提出BALOR方法,通过logit校准缓解大语言模型隐藏表示解释中的模型偏差,提升解释的忠实度和上下文信息的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22745 2026-02-02 cs.LG 57%

Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss

Softmax损失是否足够?对Softmax家族损失的系统分析

Yuanhao Pu, Defu Lian, Enhong Chen

机构 * School of Artificial Intelligence \& Data Science, University of Science \& Technology of China, Hefei, China School of Computer Science \& Technology, University of Science \& Technology of China, Hefei, China State Key Laboratory of Cognitive Intelligence, China

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本文系统分析了Softmax家族损失的理论性质和实践效果,揭示了不同替代物在分类和排序中的一致性及收敛行为,提出了偏差-方差分解和复杂度分析,为大规模类别学习中的损失选择提供了理论基础和实践指导。

Comments 34 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12767 2026-02-02 cs.AI 57%

Language Models That Walk the Talk: A Framework for Formal Fairness Certificates

语言模型言出必行:一个形式公平证书的框架

Danqing Chen, Tobias Ladner, Ahmed Rayen Mhadhbi, Matthias Althoff

机构 * Technical University of Munich, Germany(慕尼黑技术大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 本文提出一个框架,用于验证大语言模型的鲁棒性和公平性,特别是在性别公平和毒性检测中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21226 2026-01-30 cs.AI 57%

Delegation Without Living Governance

无生命治理的委托

Wolfgang Rohde

机构 * AiSuNe Foundation(AiSuNe基金会)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 本文探讨了在代理AI系统决策成为运行时决策的情况下,如何通过运行时治理(治理双胞胎)维持人类在社会、经济和政治结果塑造中的相关性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12193 2026-01-30 cs.CY 57%

The Narrow Depth and Breadth of Corporate Responsible AI Research

企业负责任的人工智能研究的深度和广度

Nur Ahmed, Amit Das, Kirsten Martin, Kawshik Banerjee

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 研究揭示企业负责任的人工智能研究存在深度和广度不足的问题,需加强公开参与以提升社会影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14401 2026-01-28 cs.MA cs.AI 57%

The Role of Social Learning and Collective Norm Formation in Fostering Cooperation in LLM Multi-Agent Systems

在LLM多智能体系统中,社会学习和集体规范形成促进合作的作用

Prateek Gupta, Qiankun Zhong, Hiromu Yakura, Thomas Eisenmann, Iyad Rahwan

机构 * Center for Humans and Machines(人类与机器中心) Max-Planck Institute for Human Development(人类发展马克斯·普朗克研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种无显式奖励信号的CPR模拟框架,通过社会学习和规范惩罚机制研究LLM多智能体系统中合作与规范的内生形成。

Comments Accepted at the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17055 2026-01-27 cs.CY 57%

AI, Metacognition, and the Verification Bottleneck: A Three-Wave Longitudinal Study of Human Problem-Solving

人工智能、元认知与验证瓶颈:人类问题解决的三波纵向研究

Matthias Huemmer, Franziska Durner, Theophile Shyiramunda, Michelle J. Cummings-Koether

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本研究探讨了生成式AI对人类问题解决的影响,发现验证成为瓶颈,提出ACTIVE框架以应对认知负荷问题。

Comments 62 pages, 2 figures, 23 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11369 2026-01-21 cs.GT cs.AI 57%

Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs

机构AI:通过公共治理图治理多智能体Cournot市场中的LLM合谋

Marcantonio Bracale Syrnikov, Federico Pierucci, Marcello Galisai, Matteo Prandi, Piercosma Bisconti, Francesco Giarrusso, Olga Sorokoletova, Vincenzo Suriani, Daniele Nardi

机构 * DEXAI – Icaro Lab(DEXAI–Icaro实验室) Sapienza University of Rome(罗马大学萨皮恩扎分校) Sant’Anna School of Advanced Studies(圣安娜高级研究学校) VU Amsterdam(阿姆斯特丹自由大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出通过公共治理图治理多智能体Cournot市场合谋问题,展示机构AI框架在减少合谋行为方面的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12727 2026-01-21 cs.HC cs.AI 57%

AI-exhibited Personality Traits Can Shape Human Self-concept through Conversations

基于AI表现的人格特质可通过对话影响人类自我概念

Jingshu Li, Tianqi Song, Nattapat Boonprakong, Zicheng Zhu, Yitian Yang, Yi-Chieh Lee

机构 * Computer Science(计算机科学) National University of Singapore(新加坡国立大学) School of Computing(计算学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本研究发现基于AI的人格特质可通过对话影响用户自我概念,揭示了AI在人机交互中的潜在影响及设计启示。

Comments ACM CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09478 2026-01-21 cs.IR cs.AI 57%

Bridging Semantic Understanding and Popularity Bias with LLMs

通过大语言模型弥合语义理解和流行偏见之间的鸿沟

Renqiang Luo, Dong Zhang, Yupeng Gao, Wen Shi, Mingliang Hou, Jiaying Liu, Zhe Wang, Shuo Yu

机构 * Jilin University Changchun China Dalian University of Technology Dalian China Jinan University \& TAL Education Group Guangzhou China Jilin University Dalian University of Technology Jinan University \& TAL Education Group

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出FairLRM框架,通过大语言模型增强对流行偏见的语义理解,提升推荐系统的公平性和准确性。

Comments 10 pages, 4 figs, WWW 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11953 2026-01-21 cs.LG 57%

Controlling Underestimation Bias in Constrained Reinforcement Learning for Safe Exploration

在安全探索中约束强化学习中的低估偏差控制

Shiqing Gao, Jiaxin Ding, Luoyi Fu, Xinbing Wang

机构 * Shanghai Jiao Tong University, Shanghai, China(上海交通大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

AI总结 本文提出MICE方法,通过引入内在成本和偏差校正策略,有效控制约束强化学习中的低估偏差,减少约束违反并保持策略性能。

Comments Published in the 42nd International Conference on Machine Learning (ICML 2025, Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11576 2026-01-21 cs.CY 57%

What Can Student-AI Dialogues Tell Us About Students' Self-Regulated Learning? An exploratory framework

学生-人工智能对话能告诉我们什么?关于学生自主学习能力的探索性框架

Long Zhang, Fangwei Lin, Weilin Wang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本研究通过分析学生与AI对话日志,提出DHASRL框架,揭示主动对话模式与自主学习能力正相关,而反应性模式则负相关。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15639 2026-01-21 cs.AI 57%

The AI Policy Module: Developing Computer Science Student Competency in AI Ethics and Policy

AI政策模块:培养计算机科学学生在AI伦理与政策方面的能力

James Weichert, Daniel Dunlap, Mohammed Farghally, Hoda Eldardiry

机构 * Computer Science \& Engineering University of Washington Seattle, USA

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出AI政策模块2.0,通过课程改革提升学生在AI伦理与政策方面的素养,通过试点评估显示学生对AI伦理影响的担忧增加,同时增强了讨论AI监管的能力。

Comments Accepted at IEEE Frontiers in Education (FIE) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10983 2026-01-19 cs.CY 57%

Evaluating 21st-Century Competencies in Postsecondary Curricula with Large Language Models: Performance Benchmarking and Reasoning-Based Prompting Strategies

利用大语言模型评估21世纪能力在高等教育课程中的表现:性能基准测试与基于推理的提示策略

Zhen Xu, Xin Guan, Chenxi Shi, Qinhao Chen, Renzhe Yu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本研究利用大语言模型评估21世纪能力在高等教育课程中的表现,提出基于推理的提示策略提升课程分析效果。

详情

展开后加载摘要…

URL PDF HTML 收藏