arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2601.00816 2026-01-06 cs.AI cs.CR cs.LG 62%

MathLedger: A Verifiable Learning Substrate with Ledger-Attested Feedback

MathLedger: 一种具有账本证明反馈的可验证学习基础

Ismail Ahmad Abdullah

机构 * CNU(中国矿业大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

AI总结 MathLedger通过整合形式验证、密码学证明和学习动态,提供一种可验证学习的基础,实现可审计的机器认知系统。

Comments 14 pages, 1 figure, 2 tables, 2 appendices with full proofs. Documents v0.9.4-pilot-audit-hardened audit surface with fail-closed governance, canonical JSON hashing, and artifact classification. Phase I infrastructure validation; no capability claims

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22725 2025-12-30 cs.CL cs.CY 62%

Mitigating Social Desirability Bias in Random Silicon Sampling

在随机硅采样中缓解社会可取性偏差

Sashank Chapala, Maksym Mironov, Songgaojun Deng

机构 * Eindhoven University of Technology(埃因霍温理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 本研究通过提示工程方法减轻LLM中的社会可取性偏差,提升硅样本与人类数据的一致性。

Comments 31 pages, 9 figures, and 24 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21775 2025-12-29 cs.AI cs.CY cs.DB 62%

Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets

合规评分方案:一种面向生成式人工智能数据集的数据溯源框架

Matyas Bohacek, Ignacio Vilanova Echavarri

机构 * Stanford University(斯坦福大学) Imperial College London(帝国理工学院伦敦分校)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文提出合规评分方案,用于评估生成式人工智能数据集的透明度、问责制和安全性,同时提供开源库以实现数据溯源,促进负责任的数据集构建与管理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13540 2025-12-29 cs.LG cs.CY 62%

Fairness-Aware Graph Representation Learning with Limited Demographic Information

带有有限人口信息的公平图表示学习

Zichong Wang, Zhipeng Yin, Liping Yang, Jun Zhuang, Rui Yu, Qingzhao Kong, Wenbin Zhang

机构 * Florida International University(佛罗里达国际大学) University of New Mexico(新墨西哥大学) Boise State University(博伊西州立大学) University of Louisville(路易斯维尔大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

AI总结 本文提出FairGLite框架,在有限人口信息下减轻图学习中的偏见,通过生成人口信息代理和自适应置信度策略,实现公平性和效用的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00931 2025-12-19 cs.CL cs.AI 62%

Mitigating Hallucinations in Zero-Shot Scientific Summarisation: A Pilot Study

缓解零样本科学摘要中的幻觉:一项初步研究

Imane Jaaouine, Ross D. King

机构 * Department of Chemical Engineering and Biotechnology(化学工程与生物技术系) University of Cambridge(剑桥大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究探讨提示工程方法对缓解零样本科学摘要中的幻觉效果,发现上下文重复和随机添加能显著提升摘要与原文的对齐度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13702 2025-12-17 cs.CY cs.AI 62%

Enhancing Transparency and Traceability in Healthcare AI: The AI Product Passport

提升医疗AI的透明度与可追溯性:AI产品护照

A. Anil Sinaci, Senan Postaci, Dogukan Cavdaroglu, Machteld J. Boonstra, Okan Mercan, Kerem Yilmaz, Gokce B. Laleci Erturkmen, Folkert W. Asselbergs, Karim Lekadir

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究提出AI产品护照,通过生命周期文档提升医疗AI的透明度和可追溯性,符合FUTURE-AI原则,确保公平性和可用性,并通过开源平台实现可访问性。

Comments A total of 33 pages: First 16 pages for the manuscript and the remaining 17 pages for the supplementary user guide of the graphical user interface

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16069 2025-12-12 cs.CY cs.AI 62%

Human or AI? Comparing Design Thinking Assessments by Teaching Assistants and Bots

人类还是人工智能?教学助教与机器人的设计思维评估比较

Sumbul Khan, Wei Ting Liow, Lay Kee Ang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究比较了教学助教与AI在评估设计思维教育学生海报中的表现,发现教师更偏好助教评分,但AI在一致性与反馈效率上具有一定优势。

Comments to be published in IEEE TALE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09483 2025-12-11 cs.CL cs.CY 62%

Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines

基于大语言模型的搜索引擎与传统搜索引擎的来源覆盖与引用偏见

Peixian Zhang, Qiming Ye, Zifan Peng, Kiran Garimella, Gareth Tyson

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Rutgers University(罗格斯大学) Rutgers University New Brunswick United States(罗格斯大学新 Brunswick美国)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.CY

AI总结 本文研究了基于大语言模型的搜索引擎与传统搜索引擎在来源覆盖和引用偏见方面的差异,发现LLM-SEs在资源多样性上优于传统搜索引擎,但其可信度和中立性仍需进一步提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09458 2025-12-11 cs.AI cs.LG 62%

Architectures for Building Agentic AI

构建代理AI的架构

Sławomir Nowaczyk

机构 * Center for Applied Intelligent Systems Research, Halmstad University, Sweden(应用智能系统研究所,哈尔姆斯塔德大学,瑞典)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种构建代理AI的架构分类,通过组件化设计和显式控制循环提升系统可靠性。

Comments This is a preprint of a chapter accepted for publication in Generative and Agentic AI Reliability: Architectures, Challenges, and Trust for Autonomous Systems, published by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08978 2025-12-11 cs.CY cs.AI 62%

Institutional AI Sovereignty Through Gateway Architecture: Implementation Report from Fontys ICT

通过网关架构实现机构AI主权:Fontys ICT的实施报告

Ruud Huijts, Koen Suilen

机构 * Fontys ICT

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出通过网关架构实现机构AI主权,构建了受监管的AI平台,实现可控的AI访问和治理,强调AI作为战略工具需专门领导和治理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08592 2025-12-10 cs.AI cs.CY cs.HC cs.SY eess.SY 62%

The SMART+ Framework for AI Systems

为AI系统设计的SMART+框架

Laxmiraju Kandikatla, Branislav Radeljic

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 SMART+框架为AI系统提供了一种结构化模型,涵盖安全、监控、责任、可靠性和透明性,并增强隐私与安全、数据治理和公平性,以提升AI系统的治理和合规性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04691 2025-12-05 cs.AI cs.CL cs.MA 62%

Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective

迈向伦理化的大型语言模型多智能体系统:从机制可解释性视角出发

Jae Hee Lee, Anne Lauscher, Stefano V. Albrecht

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文从机制可解释性视角出发,提出确保大型语言模型多智能体系统伦理行为的研究议程,聚焦于伦理评估框架、内部机制解析及参数高效对齐技术。

Comments Accepted to LaMAS 2026@AAAI'26 (https://sites.google.com/view/lamas2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23733 2025-12-04 cs.CY cs.AI cs.HC 62%

Unintentional Consequences: Generative AI Use for Cybercrime

无意后果:生成式AI用于网络犯罪

Truong Jack Luu, Binny M. Samuel

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 生成式AI的普及导致网络犯罪激增,研究通过分析数据揭示AI技术放大恶意行为的机制,并提出多层策略以应对风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16148 2025-12-03 cs.CY cs.AI 62%

Towards responsible AI for education: Hybrid human-AI to confront the Elephant in the room

迈向负责任的教育AI:混合人类-人工智能以应对教育领域的关键问题

Danial Hooshyar, Gustav Šír, Yeongwook Yang, Eve Kikas, Raija Hämäläinen, Tommi Kärkkäinen, Dragan Gašević, Roger Azevedo

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

AI总结 本文探讨教育AI中的关键问题,提出神经符号AI作为解决这些问题的混合方法,以实现负责任的AI系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02265 2025-12-03 cs.LG cs.CY 62%

The Effect of Enforcing Fairness on Reshaping Explanations in Machine Learning Models

在机器学习模型中强制公平性对解释重塑的影响

Joshua Wolff Anderson, Shyam Visweswaran

机构 * Intelligent Systems Program, University of Pittsburgh(1 智能系统计划,匹兹堡大学) Department of Biomedical Informatics, University of Pittsburgh(2 生物医学信息学系,匹兹堡大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

AI总结 本研究探讨了在医疗机器学习中通过偏见缓解技术提高公平性如何影响基于Shapley的特征排名,发现公平性提升可能改变特征重要性排名,强调了在模型评估中需综合考虑准确性、公平性和可解释性。

Comments 10 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02058 2025-12-03 cs.CY cs.CL 62%

Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings

在低资源环境下,LLM生成的人设与人类感知的错位

Tabia Tanzin Prama, Christopher M. Danforth, Peter Sheridan Dodds

机构 * Computational Story Lab(计算故事实验室) Vermont Complex Systems Institute(佛罗里达复杂系统研究所) Vermont Advanced Computing Center(佛罗里达高级计算中心) Department of Mathematics and Statistics(数学与统计学系) Department of Computer Science University of Vermont(佛罗里达大学计算机科学系) Santa Fe Institute(圣菲研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 本研究发现,在低资源环境中,LLM生成的人设在共情和可信度方面显著劣于人类,需通过现实数据验证以确保其可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00742 2025-12-02 cs.CY cs.AI 62%

On the Regulatory Potential of User Interfaces for AI Agent Governance

关于用户界面在AI代理治理中的调节潜力

K. J. Kevin Feng, Tae Soo Kim, Rock Yuren Pang, Faria Huq, Tal August, Amy X. Zhang

机构 * University of Washington(华盛顿大学) KAIST(韩国科学技术院) Carnegie Mellon University(卡内基梅隆大学) UIUC(伊利诺伊大学香槟分校)

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.CY

AI总结 本文提出通过调节AI代理的用户界面来增强透明性和行为规范,从而在系统和基础设施层面实现治理。

Comments RegML workshop at NeurIPS 2025 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00461 2025-12-02 cs.CY cs.CL 62%

Whose Personae? Synthetic Persona Experiments in LLM Research and Pathways to Transparency

谁的人设?LLM研究中的人设实验及透明化路径

Jan Batzner, Volker Stocker, Bingjun Tang, Anusha Natarajan, Qinhao Chen, Stefan Schmid, Gjergji Kasneci

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 本文探讨了LLM研究中合成人设实验的代表性与生态效度问题,提出透明化检查表以提升评估的严谨性和实证性。

Comments Published at AAAI/ACM AIES 2025. Presented at NeurIPS 2025 Workshop Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

Journal ref Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8(1), 2025, 343-354

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20680 2025-11-27 cs.CL cs.AI 62%

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

大语言模型的认知偏差影响临床肿瘤学笔记的解读

Matthew W. Kenaston, Umair Ayub, Mihir Parmar, Muhammad Umair Anjum, Syed Arsalan Ahmed Naqvi, Priya Kumar, Samarth Rawal, Aadel A. Chaudhuri, Yousef Zakharia, Elizabeth I. Heath, Tanios S. Bekaii-Saab, Cui Tao, Eliezer M. Van Allen, Ben Zhou, YooJung Choi, Chitta Baral, Irbaz Bin Riaz

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本研究揭示了大语言模型在肿瘤学笔记解读中因推理缺陷导致的临床安全隐患,并提出了一种可推广的推理错误分类框架。

Comments 24 pages, 6 figures, 1 supplementary figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18182 2025-11-25 cs.CY cs.AI 62%

The Workflow as Medium: A Framework for Navigating Human-AI Co-Creation

流程作为媒介:一种导航人机协同创作的框架

Lee Ackerman

机构 * Media University of Applied Sciences(应用科学媒体大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出创意智能循环框架,通过图文小说探讨人工智能在人机协同创作中的伦理与治理挑战,推动AI素养提升。

Comments 57 pages, 13 images, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04695 2025-11-21 cs.AI cs.CE cs.ET cs.LG 62%

Bridging the Gap in XAI-Why Reliable Metrics Matter for Explainability and Compliance

弥合XAI差距:可靠度量在可解释性与合规性中的重要性

Pratinav Seth, Vinay Kumar Sankarapu

机构 * Lexsi Labs(Lexsi实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出以度量治理的范式,通过标准化指标提升AI系统的可解释性和合规性,防止对齐造假,构建持续的AI保证流程。

Comments Accepted at first EurIPS Workshop on Private AI Governance

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14606 2025-11-19 cs.CL cs.LG 62%

Bridging Human and Model Perspectives: A Comparative Analysis of Political Bias Detection in News Media Using Large Language Models

Shreya Adrita Banik, Niaz Nafi Rahman, Tahsina Moiukh, Farig Sadeque

机构 * Department of Computer Science and Engineering, BRAC University(计算机科学与工程系,BRAC大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18708 2025-11-19 cs.MA cs.AI cs.LG 62%

Skill-Aligned Fairness in Multi-Agent Learning for Collaboration in Healthcare

Promise Osaine Ekpo, Brian La, Thomas Wiener, Saesha Agarwal, Arshia Agrawal, Gonzalo Gonzalez-Pumariega, Lekan P. Molu, Angelique Taylor

机构 * Cornell Tech(康奈尔科技) Microsoft Research(微软研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11790 2025-11-18 cs.CY cs.AI 62%

Differences in the Moral Foundations of Large Language Models

Peter Kirgis

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10089 2025-11-18 cs.LG cs.AI 62%

T2IBias: Uncovering Societal Bias Encoded in the Latent Space of Text-to-Image Generative Models

Abu Sufian, Cosimo Distante, Marco Leo, Hanan Salam

机构 * National Research Council of Italy - Institute of Applied Sciences

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments This manuscript has been accepted for presentation in the First Interdisciplinary Workshop on Responsible AI for Value Creation. Dec 1, Copenhagen. The final version will be submitted for inclusion in a Springer LNCS Volume. (The paper is 15 pages with 7 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05927 2025-11-13 cs.CY cs.AI econ.GN q-fin.EC 62%

Artificial intelligence and the Gulf Cooperation Council workforce adapting to the future of work

Mohammad Rashed Albous, Melodena Stephens, Odeh Rashed Al-Jayyousi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Journal ref Humanit Soc Sci Commun 12, 1649 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08567 2025-11-12 cs.LG cs.AI 62%

The Path Not Taken: RLVR Provably Learns Off the Principals

Hanqing Zhu, Zhenyu Zhang, Hanxian Huang, DiJia Su, Zechun Liu, Jiawei Zhao, Igor Fedorov, Hamed Pirsiavash, Zhizhou Sha, Jinwon Lee, David Z. Pan, Zhangyang Wang, Yuandong Tian, Kai Sheng Tai

机构 * Meta AI The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments Preliminary version accepted as a spotlight in NeurIPS 2025 Workshop on Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08082 2025-11-12 cs.AI cs.LG econ.GN q-fin.EC 62%

Prudential Reliability of Large Language Models in Reinsurance: Governance, Assurance, and Capital Efficiency

Stella C. Dong

机构 * Reinsurance Analytics(再保险分析)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments 48 pages, 9 figures, 5 tables. Submitted to the Journal of Risk and Insurance (JRI), November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07803 2025-11-12 cs.CY cs.AI 62%

Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring

Wenhao Xu, Akshatha Arodi, Jian-Yun Nie, Arsene Fansi Tchango

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments To appear at AAAI-26 (Social Impact Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04698 2025-11-11 cs.CL cs.AI 62%

multiMentalRoBERTa: A Fine-tuned Multiclass Classifier for Mental Health Disorder

K M Sajjadul Islam, John Fields, Praveen Madiraju

机构 * Marquette University, WI, USA(马quette大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted in IEEE Big Data, 8-11 December, 2025 @ Macau SAR, China

详情

展开后加载摘要…

URL PDF HTML 收藏