arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2511.02895 2025-11-07 cs.CY cs.AI cs.HC physics.soc-ph 73%

A Criminology of Machines

Gian Maria Campedelli

机构 * Fondazione Bruno Kessler(布鲁诺·科斯勒基金会)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments This pre-print is also available at CrimRxiv with DOI: https://doi.org/10.21428/cb6ab371.e3354ce1

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25445 2025-10-30 cs.AI cs.LG 73%

Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions

Mohamad Abou Ali, Fadi Dornaika

机构 * University of the Basque Country(巴斯克大学) Lebanese International University (LIU)(黎巴嫩国际大学) The International University of Beirut(贝鲁特国际大学) IKERBASQUE

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19008 2025-10-23 cs.HC cs.AI cs.LG cs.MA 73%

Plural Voices, Single Agent: Towards Inclusive AI in Multi-User Domestic Spaces

Joydeep Chandra, Satyam Kumar Navneet

机构 * BNRIST, Tsinghua University(北京理工大学、清华大学) Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16555 2025-10-21 cs.AI cs.LG 73%

Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence

Qiongyan Wang, Xingchen Zou, Yutian Jiang, Haomin Wen, Jiaheng Wei, Qingsong Wen, Yuxuan Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Carnegie Mellon University(卡内基梅隆大学) Squirrel Ai Learning

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07887 2025-10-17 cs.CL cs.AI 73%

Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge

Riccardo Cantini, Alessio Orsino, Massimo Ruggiero, Domenico Talia

机构 * University of Calabria(卡利博利亚大学)

专题命中 AI治理与伦理 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Journal ref Cantini, R., Orsino, A., Ruggiero, M., Talia, D. Benchmarking adversarial robustness to bias elicitation in large language models: scalable automated assessment with LLM-as-a-judge. Mach Learn 114, 249 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02444 2025-10-08 cs.CL cs.AI 73%

Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models

Haoran Ye, Tianze Zhang, Yuhang Xie, Liyuan Zhang, Yuanyi Ren, Xin Zhang, Guojie Song

机构 * State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(人工智能通用基础理论国家重点实验室,智能科学与技术学院,北京大学) Yuanpei College, Peking University(元培学院,北京大学) School of Psychological and Cognitive Sciences, Peking University(心理学与认知科学学院,北京大学) Key Laboratory of Machine Perception (Ministry of Education), Peking University(机器感知重点实验室(教育部),北京大学) PKU-Wuhan Institute for Artificial Intelligence(北京大学-武汉人工智能研究院)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04931 2025-09-04 cs.AI cs.CL cs.HC 73%

A Survey on Human-AI Collaboration with Large Foundation Models

Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yanwen Xu, Lei Ding, Jay Mehta, Nathan Grinnell, Li Liu, Sijia Zhong, Devanathan Nallur Gandamani, Xinyi Tang, Rohan Ghosalkar, Celeste Shen, Rachel Shen, Nafisa Hussain, Kesav Ravichandran, James Davis

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments Topic and scope refinement

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00085 2025-09-03 cs.CR cs.AI cs.CY 73%

Private, Verifiable, and Auditable AI Systems

Tobin South

机构 * MIT Media Lab(麻省理工学院媒体实验室)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06849 2025-08-12 cs.CY cs.AI cs.HC 73%

Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development

Sanjana Gautam, Mohit Chandra, Ankolika De, Tatiana Chakravorti, Girik Malik, Munmun De Choudhury

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11502 2025-07-16 cs.CL cs.CE cs.LG 73%

HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong

Sirui Han, Junqi Zhu, Ruiyuan Zhang, Yike Guo

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03050 2025-07-08 cs.CY cs.AI 73%

From Turing to Tomorrow: The UK's Approach to AI Regulation

Oliver Ritchie, Markus Anderljung, Tom Rachman

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments This is a chapter intended for publication in a forthcoming edited volume. It is the version of the author's manuscript prior to acceptance for publication and has not undergone editorial and/or peer review on behalf of the Publisher

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05050 2025-06-04 cs.CL cs.AI 73%

Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models

Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Shaohui Mei, Lap-Pui Chau

机构 * Department of Electrical and Electronic Engineering(电子与电气工程系) The Hong Kong Polytechnic University(香港理工大学) School of Electronics and Information(电子与信息学院) Northwestern Polytechnical University(西北工业大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01257 2025-06-03 cs.CL cs.AI 73%

DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models

Jiancheng Ye, Sophie Bronstein, Jiarui Hai, Malak Abu Hashish

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17131 2025-05-26 cs.CL cs.AI stat.ML 73%

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs

Alireza Arbabi, Florian Kerschbaum

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04756 2025-05-23 cs.CL cs.LG 73%

Red-Teaming for Inducing Societal Bias in Large Language Models

Chu Fei Luo, Ahmad Ghawanmeh, Bharat Bhimshetty, Kashyap Murali, Murli Jadhav, Xiaodan Zhu, Faiza Khan Khattak

机构 * Queen’s University Vector Institute(女王大学向量研究所) SigmaRed Tech.(SigmaRed科技公司) Ernst & Young(埃森哲) Monark Health(Monark健康)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13972 2025-04-22 cs.CY cs.AI 73%

Governance Challenges in Reinforcement Learning from Human Feedback: Evaluator Rationality and Reinforcement Stability

Dana Alsagheer, Abdulrahman Kamal, Mohammad Kamal, Weidong Shi

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09097 2024-12-18 cs.CL cs.AI 73%

Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations

Tarun Raheja, Nilay Pochhi, F. D. C. M. Curie

专题命中 AI治理与伦理 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Comments 16 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14194 2024-10-21 cs.CL cs.AI 73%

Speciesism in Natural Language Processing Research

Masashi Takeshita, Rafal Rzepka

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments This article is a preprint and has not been peer-reviewed. The postprint has been accepted for publication in AI and Ethics. Please cite the final version of the article once it is published

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08323 2024-09-18 cs.CY cs.AI 73%

Mapping the Ethics of Generative AI: A Comprehensive Scoping Review

Thilo Hagendorff

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04023 2024-08-09 cs.CL cs.AI 73%

Improving Large Language Model (LLM) fidelity through context-aware grounding: A systematic approach to reliability and veracity

Wrick Talukdar, Anjanava Biswas

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments 14 pages

Journal ref World Journal of Advanced Engineering Technology and Sciences, 2023, 10(2), 283-296

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03820 2024-06-14 cs.CY cs.AI cs.HC 73%

False Sense of Security in Explainable Artificial Intelligence (XAI)

Neo Christopher Chung, Hongkyou Chung, Hearim Lee, Lennart Brocki, Hongbeom Chung, George Dyer

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

Comments AI Governance Workshop at the 2024 International Joint Conference on Artificial Intelligence (IJCAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12744 2024-05-13 cs.CL cs.AI 73%

Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches

Pablo Biedma, Xiaoyuan Yi, Linus Huang, Maosong Sun, Xing Xie

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments 16 pages, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01886 2024-05-06 cs.CL cs.AI 73%

Aloe: A Family of Fine-tuned Open Healthcare LLMs

Ashwin Kumar Gururajan, Enrique Lopez-Cuena, Jordi Bayarri-Planas, Adrian Tormos, Daniel Hinjos, Pablo Bernabeu-Perez, Anna Arias-Duart, Pablo Agustin Martin-Torres, Lucia Urcelay-Ganzabal, Marta Gonzalez-Mallo, Sergio Alvarez-Napagao, Eduard Ayguadé-Parra, Ulises Cortés Dario Garcia-Gasulla

专题命中 AI治理与伦理 :alignment(abstract);red teaming(abstract);分类 cs.CL、cs.AI

Comments Five appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19076 2024-05-01 cs.CY cs.AI 73%

Who Followed the Blueprint? Analyzing the Responses of U.S. Federal Agencies to the Blueprint for an AI Bill of Rights

Darren Lage, Riley Pruitt, Jason Ross Arnold

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06263 2023-11-14 cs.CY cs.AI 73%

No Trust without regulation!

François Terrier

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15370 2023-06-07 cs.CY cs.AI 73%

A Principles-based Ethics Assurance Argument Pattern for AI and Autonomous Systems

Zoe Porter, Ibrahim Habli, John McDermid, Marten Kaas

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

Journal ref AI and Ethics 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11163 2023-04-25 cs.CY cs.CL 73%

ChatGPT, Large Language Technologies, and the Bumpy Road of Benefiting Humanity

Atoosa Kasirzadeh

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.CY

Comments As part of a series on Dailynous : "Philosophers on next-generation large language models"

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04776 2023-01-04 cs.LG cs.AI cs.CV cs.HC 73%

What should AI see? Using the Public's Opinion to Determine the Perception of an AI

Robin Chan, Radin Dardashti, Meike Osinski, Matthias Rottmann, Dominik Brüggemann, Cilia Rücker, Peter Schlicht, Fabian Hüger, Nikol Rummel, Hanno Gottschalk

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments 26 pages, 12 figures

Journal ref AI and Ethics (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.04520 2022-11-30 cs.LG cs.AI 73%

Obtaining Dyadic Fairness by Optimal Transport

Moyi Yang, Junjie Sheng, Xiangfeng Wang, Wenyan Liu, Bo Jin, Jun Wang, Hongyuan Zha

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11446 2022-01-24 cs.CL cs.AI 73%

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d'Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, Geoffrey Irving

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments 120 pages

详情

展开后加载摘要…

URL PDF HTML 收藏