arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2510.06105 2025-10-08 cs.AI cs.CY cs.HC cs.LG 67%

Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences

Batu El, James Zou

机构 * Stanford University(斯坦福大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10127 2025-10-07 cs.CL cs.AI cs.LG 67%

Population-Aligned Persona Generation for LLM-based Social Simulation

Zhengyu Hu, Jianxun Lian, Zheyuan Xiao, Max Xiong, Yuxuan Lei, Tianfu Wang, Kaize Ding, Ziang Xiao, Nicholas Jing Yuan, Xing Xie

机构 * HKUST(香港科技大学) Microsoft Research Asia(微软亚洲研究院) Duke University(杜克大学) Northwestern University(西北大学) Johns Hopkins University(约翰霍普金斯大学) Microsoft(微软)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01164 2025-10-02 cs.CL cs.AI cs.CY cs.HC 67%

Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare

Zhengliang Shi, Ruotian Ma, Jen-tse Huang, Xinbei Ma, Xingyu Chen, Mengru Wang, Qu Yang, Yue Wang, Fanghua Ye, Ziyang Chen, Shanyi Wang, Cixing Li, Wenxuan Wang, Zhaopeng Tu, Xiaolong Li, Zhaochun Ren, Linus

机构 * Tencent(腾讯)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03699 2025-09-23 cs.AI cs.CY cs.LG cs.MA q-bio.QM 67%

Enhancing Clinical Decision-Making: Integrating Multi-Agent Systems with Ethical AI Governance

Ying-Jung Chen, Ahmad Albarqawi, Chi-Sheng Chen

机构 * College of Computing(计算学院) Georgia Institute of Technology(佐治亚理工学院) University of Illinois(伊利诺伊大学) Neuro Industry Research(神经产业研究) Neuro Industry, Inc.(神经产业公司)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11648 2025-09-16 cs.CL cs.AI cs.CY 67%

EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI

Sai Kartheek Reddy Kasu

机构 * IIIT Dharwad(德瓦德理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17044 2025-09-16 cs.CY cs.AI cs.LG 67%

Approaches to Responsible Governance of GenAI in Organizations

Dhari Gandhi, Himanshu Joshi, Lucas Hartman, Shabnam Hassani

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) Western University(西部大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03034 2025-09-04 cs.LG cs.AI cs.CR cs.CV cs.CY 67%

Rethinking Data Protection in the (Generative) Artificial Intelligence Era

Yiming Li, Shuo Shao, Yu He, Junfeng Guo, Tianwei Zhang, Zhan Qin, Pin-Yu Chen, Michael Backes, Philip Torr, Dacheng Tao, Kui Ren

机构 * The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) Nanyang Technological University(南洋理工大学) University of Maryland(马里兰大学) IBM Research(IBM研究院) CISPA Helmholtz Center for Information Security(CISPA 欧洲信息安全部分) University of Oxford(牛津大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Perspective paper for a broader scientific audience. The first two authors contributed equally to this paper. 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19269 2025-08-28 cs.CY cs.AI cs.CL 67%

Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models

Ke Zhou, Marios Constantinides, Daniele Quercia

机构 * Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments This paper has been accepted in AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15860 2025-08-22 cs.CL cs.AI cs.LG 67%

Synthetic vs. Gold: The Role of LLM Generated Labels and Data in Cyberbullying Detection

Arefeh Kazemi, Sri Balaaji Natarajan Kalaivendan, Joachim Wagner, Hamza Qadeer, Kanishk Verma, Brian Davis

机构 * School of Computing, ADAPT Centre, Dublin City University, Dublin, Ireland(计算学院、ADAPT中心、都柏林城市大学、都柏林、爱尔兰)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07284 2025-08-12 cs.CL cs.AI cs.CY 67%

"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas

Junchen Ding, Penghao Jiang, Zihao Xu, Ziqi Ding, Yichen Zhu, Jiaojiao Jiang, Yuekang Li

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16663 2025-08-12 cs.LG cs.AI cs.CY cs.LO cs.SE 67%

Runtime Monitoring and Enforcement of Conditional Fairness in Generative AIs

Chih-Hong Cheng, Changshun Wu, Xingyu Zhao, Saddek Bensalem, Harald Ruess

机构 * Chalmers University of Technology, Sweden Carl von Ossietzky Universität Oldenburg, Germany Universit\'e Grenoble Alpes, France University of Warwick, United Kingdom CSX-AI, France SRI International, United States

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14339 2025-07-22 cs.CY cs.AI cs.HC cs.LG eess.SP 67%

Fiduciary AI for the Future of Brain-Technology Interactions

Abhishek Bhattacharjee, Jack Pilkington, Nita Farahany

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05442 2025-07-16 cs.AI cs.CY cs.HC cs.LG 67%

The Odyssey of the Fittest: Can Agents Survive and Still Be Good?

Dylan Waldner, Risto Miikkulainen

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted to CogSci 2025. Code can be found at https://github.com/dylanwaldner/BeGoodOrSurvive

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06434 2025-07-10 cs.CY cs.AI cs.LG 67%

Deprecating Benchmarks: Criteria and Framework

Ayrton San Joaquin, Rokas Gipiškis, Leon Staufer, Ariel Gil

机构 * AI Standards Lab(AI标准实验室) Technical University of Munich(慕尼黑技术大学) Institute of Data Science(数据科学研究所) Digital Technologies(数字技术) Trajectory Labs(轨迹实验室)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 10 pages, 1 table. Accepted to the ICML 2025 Technical AI Governance Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01630 2025-06-19 cs.LG cs.AI cs.CY 67%

Machine Learners Should Acknowledge the Legal Implications of Large Language Models as Personal Data

Henrik Nolte, Michèle Finck, Kristof Meding

机构 * University Tübingen(图宾根大学) CZS Institute for Artificial Intelligence and Law(法律与人工智能研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14230 2025-06-12 cs.CL cs.AI cs.CY 67%

Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing

Han Jiang, Xiaoyuan Yi, Zhihua Wei, Ziang Xiao, Shu Wang, Xing Xie

机构 * Tongji University(同济大学) Johns Hopkins University(约翰霍普金斯大学) University of California, Los Angeles(加州大学洛杉矶分校) Microsoft Research Asia(微软亚洲研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07045 2025-06-05 cs.CR cs.AI cs.CL cs.CY 67%

Scalable and Ethical Insider Threat Detection through Data Synthesis and Analysis by LLMs

Haywood Gelman, John D. Hastings

机构 * The Beacom College of Computer and Cyber Sciences(贝科姆计算机与网络科学学院) Dakota State University(达科他州立大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 6 pages, 0 figures, 8 tables

Journal ref 2025 IEEE 13th International Symposium on Digital Forensics and Security (ISDFS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21634 2025-05-01 cs.CY cs.AI cs.LG 67%

Quantitative Auditing of AI Fairness with Differentially Private Synthetic Data

Chih-Cheng Rex Yuan, Bow-Yaw Wang

机构 * Institute of Information Science, Academia Sinica(学术院信息研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05210 2025-04-08 cs.CY cs.AI cs.HC cs.LG 67%

A moving target in AI-assisted decision-making: Dataset shift, model updating, and the problem of update opacity

Joshua Hatherley

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Journal ref Ethics and Information Technology 27(2): 20 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15514 2025-04-08 cs.HC cs.AI cs.CL cs.CY cs.ET 67%

Superhuman Game AI Disclosure: Expertise and Context Moderate Effects on Trust and Fairness

Jaymari Chua, Chen Wang, Lina Yao

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05315 2025-02-25 cs.CL cs.AI cs.LG 67%

Aligned at the Start: Conceptual Groupings in LLM Embeddings

Mehrdad Khatir, Sanchit Kabra, Chandan K. Reddy

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14143 2025-02-21 cs.MA cs.AI cs.CY cs.ET cs.LG 67%

Multi-Agent Risks from Advanced AI

Lewis Hammond, Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, Chandler Smith, Wolfram Barfuss, Jakob Foerster, Tomáš Gavenčiak, The Anh Han, Edward Hughes, Vojtěch Kovařík, Jan Kulveit, Joel Z. Leibo, Caspar Oesterheld, Christian Schroeder de Witt, Nisarg Shah, Michael Wellman, Paolo Bova, Theodor Cimpeanu, Carson Ezell, Quentin Feuillade-Montixi, Matija Franklin, Esben Kran, Igor Krawczuk, Max Lamparth, Niklas Lauffer, Alexander Meinke, Sumeet Motwani, Anka Reuel, Vincent Conitzer, Michael Dennis, Iason Gabriel, Adam Gleave, Gillian Hadfield, Nika Haghtalab, Atoosa Kasirzadeh, Sébastien Krier, Kate Larson, Joel Lehman, David C. Parkes, Georgios Piliouras, Iyad Rahwan

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Cooperative AI Foundation, Technical Report #1

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10868 2025-01-31 cs.AI cs.CL cs.CY cs.HC 67%

From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape

Timothy R. McIntosh, Teo Susnjak, Tong Liu, Paul Watters, Malka N. Halgamuge

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16451 2024-12-24 cs.LG cs.AI cs.CL 67%

Correcting Large Language Model Behavior via Influence Function

Han Zhang, Zhuo Zhang, Yi Zhang, Yuanzhao Zhai, Hanyang Peng, Yu Lei, Yue Yu, Hui Wang, Bin Liang, Lin Gui, Ruifeng Xu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10939 2024-12-17 cs.CY cs.AI cs.CL cs.HC 67%

Human-Centric NLP or AI-Centric Illusion?: A Critical Investigation

Piyapath T Spencer

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Preprint to be published in Proceedings of PACLIC38

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00067 2024-11-25 cs.AI cs.CY cs.LG 67%

Transformers in Healthcare: A Survey

Subhash Nerella, Sabyasachi Bandyopadhyay, Jiaqing Zhang, Miguel Contreras, Scott Siegel, Aysegul Bumin, Brandon Silva, Jessica Sena, Benjamin Shickel, Azra Bihorac, Kia Khezeli, Parisa Rashidi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Journal ref Transformers and large language models in healthcare: A review, Artificial Intelligence in Medicine, Volume 154, 2024, 102900,

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17459 2024-10-24 cs.LG cs.AI cs.CR cs.CY 67%

Data Obfuscation through Latent Space Projection (LSP) for Privacy-Preserving AI Governance: Case Studies in Medical Diagnosis and Finance Fraud Detection

Mahesh Vaijainthymala Krishnamoorthy

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 19 pages, 6 figures, submitted to Conference ICADCML2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12142 2024-10-23 cs.LG cs.AI cs.CV cs.CY 67%

Slicing Through Bias: Explaining Performance Gaps in Medical Image Analysis using Slice Discovery Methods

Vincent Olesen, Nina Weng, Aasa Feragen, Eike Petersen

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments MICCAI 2024 Workshop on Fairness of AI in Medical Imaging

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13784 2024-10-21 cs.LG cs.AI cs.CY cs.SE 67%

The Model Openness Framework: Promoting Completeness and Openness for Reproducibility, Transparency, and Usability in Artificial Intelligence

Matt White, Ibrahim Haddad, Cailean Osborne, Xiao-Yang Yanglet Liu, Ahmed Abdelmonsef, Sachin Varghese, Arnaud Le Hors

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 28 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12700 2024-10-17 cs.CV cs.AI cs.CY cs.LG cs.MM 67%

Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization

Xingqi Wang, Xiaoyuan Yi, Xing Xie, Jia Jia

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted by ACM Multimedia 2024. The dataset and code can be found at https://github.com/achernarwang/LiVO

详情

展开后加载摘要…

URL PDF HTML 收藏