arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2509.05627 2025-09-09 cs.CY cs.LG stat.ML 62%

Audits Under Resource, Data, and Access Constraints: Scaling Laws For Less Discriminatory Alternatives

Sarah H. Cen, Salil Goyal, Zaynah Javed, Ananya Karthik, Percy Liang, Daniel E. Ho

机构 * Stanford University(斯坦福大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

Comments 34 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08193 2025-09-05 cs.CY cs.AI 62%

Street-Level AI: Are Large Language Models Ready for Real-World Judgments?

Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, Sanmay Das

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments This work has been accepted for publication as a full paper at the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01576 2025-09-03 cs.AI cs.CY cs.SY eess.SY 62%

Structured AI Decision-Making in Disaster Management

Julian Gerald Dcruz, Argyrios Zolotas, Niall Ross Greenwood, Miguel Arana-Catania

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 40 pages, 14 figures, 16 tables. To be published in Nature Scientific Reports

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21788 2025-09-01 cs.CL cs.AI cs.IR 62%

Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval

Inés Altemir Marinas, Anastasiia Kucherenko, Andrei Kucharavy

机构 * École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院) Institute of Entrepreneurship and Management, HES-SO Valais-Wallis(创业与管理研究所) Institute of Informatics, HES-SO Valais-Wallis(信息研究所)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20015 2025-08-28 cs.LG cs.AI 62%

Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment

Julian Arnold, Niels Lörch

机构 * Department of Physics University of Basel(物理系 巴塞尔大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments 11+25 pages, 4+11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16762 2025-08-26 cs.CL cs.CY 62%

Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation

Arka Mukherjee, Shreya Ghosh

机构 * KIIT Deemed University(KIIT大学) Indian Institute of Technology (IIT) Bhubaneswar(印度理工学院(Bhubaneswar分校))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

Comments Accepted at ASI @ ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14116 2025-08-21 cs.CY cs.AI 62%

Enriching Moral Perspectives on AI: Concepts of Trust amongst Africans

Lameck Mbangula Amugongo, Nicola J Bidwell, Joseph Mwatukange

机构 * Namibia University of Science \& Technology 13 Jackson Kaujeua Windhoek Namibia 9000 Rhodes University Makhanda South Africa International University of Management Namibia Charles Darwin University Australia Namibia University of Science \& Technology Rhodes University International University of Management Charles Darwin University

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08804 2025-08-13 cs.LG cs.AI 62%

TechOps: Technical Documentation Templates for the AI Act

Laura Lucaj, Alex Loosley, Hakan Jonsson, Urs Gasser, Patrick van der Smagt

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08544 2025-08-13 cs.CY cs.AI 62%

AI Agents and the Law

Mark O. Riedl, Deven R. Desai

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 2025 AAAI Conference on AI, Ethics, and Society

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05269 2025-08-13 cs.LG cs.AI q-bio.QM 62%

Chemist-aligned retrosynthesis by ensembling diverse inductive bias models

Krzysztof Maziarz, Guoqing Liu, Hubert Misztela, Austin Tripp, Junren Li, Aleksei Kornev, Piotr Gaiński, Holger Hoefling, Mike Fortunato, Rishi Gupta, Marwin Segler

机构 * Microsoft Research AI for Science(微软研究院人工智能与科学研究中心) Novartis Biomedical Research(诺华生物医学研究) University of Cambridge(剑桥大学) Jagiellonian University(雅盖隆大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08333 2025-08-13 cs.CY cs.AI 62%

Normative Moral Pluralism for AI: A Framework for Deliberation in Complex Moral Contexts

David-Doron Yaacov

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Conference version: AIES 2025 (non-archival track), 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07673 2025-08-12 cs.AI cs.LG 62%

Ethics2vec: aligning automatic agents and human preferences

Gianluca Bontempi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07111 2025-08-12 cs.CL cs.AI 62%

Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution

Falaah Arif Khan, Nivedha Sivakumar, Yinong Oliver Wang, Katherine Metcalf, Cezanne Camacho, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff

机构 * Apple(苹果公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05913 2025-08-11 cs.HC cs.AI cs.CL 62%

Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction

Stefan Pasch, Min Chul Cha

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03970 2025-08-07 cs.CL cs.AI 62%

Data and AI governance: Promoting equity, ethics, and fairness in large language models

Alok Abhishek, Lisa Erickson, Tushar Bandopadhyay

机构 * Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments Published in MIT Science Policy Review 6, 139-146 (2025)

Journal ref MIT Science Policy Review, 6. (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03292 2025-08-06 cs.CL cs.AI 62%

Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes

Shahed Masoudian, Gustavo Escobedo, Hannah Strauss, Markus Schedl

机构 * Johannes Kepler University (JKU)(约翰内斯·开普勒大学) Linz Institute of Technology (LIT)(林茨技术研究所) University of Innsbruck(因斯布鲁克大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09053 2025-08-06 cs.AI cs.GT cs.LG 62%

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers

Haoran Sun, Yusen Wu, Peng Wang, Wei Chen, Yukun Cheng, Xiaotie Deng, Xu Chu

机构 * CFCS, School of Computer Science, Peking University(计算机科学系,北京大学) School of Business, Jiangnan University(商学院,江南大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments A shorter conference version is published in IJCAI 2025, titled 'Game Theory Meets Large Language Models: A Systematic Survey'

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06072 2025-08-04 cs.CL cs.AI 62%

A Survey on Post-training of Large Language Models

Guiyao Tie, Zeli Zhao, Dingjie Song, Fuyang Wei, Rong Zhou, Yurou Dai, Wen Yin, Zhejian Yang, Jiangyue Yan, Yao Su, Zhenhan Dai, Yifeng Xie, Yihan Cao, Lichao Sun, Pan Zhou, Lifang He, Hechang Chen, Yu Zhang, Qingsong Wen, Tianming Liu, Neil Zhenqiang Gong, Jiliang Tang, Caiming Xiong, Heng Ji, Philip S. Yu, Jianfeng Gao

机构 * Huazhong University of Science and Technology(华中科技大学) Lehigh University(莱斯大学) The University of Hong Kong(香港大学) Jilin University(吉林大学) Southern University of Science and Technology(南方科技大学) Worcester Polytechnic Institute(沃思堡理工学院) LinkedIn Corporation(领英公司) Squirrel Ai Learning University of Georgia(佐治亚大学) Duke University(杜克大学) Michigan State University(密歇根州立大学) Salesforce Research(Salesforce研究) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Illinois at Chicago(伊利诺伊大学芝加哥分校) Microsoft Research(微软研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 87 pages, 21 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22445 2025-07-31 cs.CL cs.AI 62%

AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini

Jill Walker Rettberg, Hermann Wigers

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments This project has received funding from the European Union's Horizon 2020 research and innovation programme under grant agreement number 101142306. The project is also supported by the Center for Digital Narrative, which is funded by the Research Council of Norway through its Centres of Excellence scheme, project number 332643

Journal ref Open Research Europe 2025, 5:202 [version 1; peer review: awaiting peer review]

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21083 2025-07-30 cs.CL cs.AI 62%

ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs

Franck Bardol

机构 * Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17788 2025-07-25 cs.LG cs.AI 62%

Adaptive Repetition for Mitigating Position Bias in LLM-Based Ranking

Ali Vardasbi, Gustavo Penha, Claudia Hauff, Hugues Bouchard

机构 * Spotify

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17787 2025-07-25 cs.LG cs.AI 62%

Hyperbolic Deep Learning for Foundation Models: A Survey

Neil He, Hiren Madhu, Ngoc Bui, Menglin Yang, Rex Ying

机构 * Yale University(耶鲁大学) Hong Kong University of Science(香港科学大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments 11 Pages, SIGKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15901 2025-07-23 cs.AI cs.CY cs.MA 62%

Advancing Responsible Innovation in Agentic AI: A study of Ethical Frameworks for Household Automation

Joydeep Chandra, Satyam Kumar Navneet

机构 * Department of CST Tsinghua University Beijing, China(计算机科学与技术系 清华大学 北京中国) Department of CSE Chandigarh University Mohali, India(计算机科学与工程系 印度昌迪加尔大学 摩哈利)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15885 2025-07-23 cs.AI cs.HC cs.LG 62%

ADEPTS: A Capability Framework for Human-Centered Agent Design

Pierluca D'Oro, Caley Drooff, Joy Chen, Joseph Tighe

机构 * FAIR at Meta(Meta 的 FAIR)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11415 2025-07-15 cs.CL cs.AI 62%

Political Bias in LLMs: Unaligned Moral Values in Agent-centric Simulations

Simon Münker

机构 * Journal for Language Technology and Computational Linguistics(语言技术与计算语言学期刊)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 2 tables

Journal ref JLCL 2025, Band 38(2)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08908 2025-07-15 cs.CY cs.ET cs.LG 62%

The Engineer's Dilemma: A Review of Establishing a Legal Framework for Integrating Machine Learning in Construction by Navigating Precedents and Industry Expectations

M. Z. Naser

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05443 2025-07-03 cs.CL cs.AI 62%

A survey of textual cyber abuse detection using cutting-edge language models and large language models

Jose A. Diaz-Garcia, Joao Paulo Carvalho

机构 * Department of Computer Science and A.I, University of Granada(计算机科学与人工智能系,格拉纳达大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments 37 pages, under review in WIREs Data Mining and Knowledge Discovery

Journal ref Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery (2025), 15(3), e70029

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15620 2025-06-19 cs.LG cs.AI 62%

GFLC: Graph-based Fairness-aware Label Correction for Fair Classification

Modar Sulaiman, Kallol Roy

机构 * University of Tartu, Institute of Computer Science(塔尔图大学计算机科学研究所)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 25 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14645 2025-06-18 cs.CL cs.CY 62%

Passing the Turing Test in Political Discourse: Fine-Tuning LLMs to Mimic Polarized Social Media Comments

. Pazzaglia, V. Vendetti, L. D. Comencini, F. Deriu, V. Modugno

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11627 2025-06-16 cs.CV cs.AI cs.LG 62%

Evaluating Fairness and Mitigating Bias in Machine Learning: A Novel Technique using Tensor Data and Bayesian Regression

Kuniko Paxton, Koorosh Aslansefat, Dhavalkumar Thakker, Yiannis Papadopoulos

机构 * School of Computer Science and DAIM University of Hull(计算机科学学院和DAIM赫尔大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏