arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1832 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1832 篇

2505.13673 2025-05-21 cs.CY 57%

Comparing Apples to Oranges: A Taxonomy for Navigating the Global Landscape of AI Regulation

Sacha Alanoca, Shira Gur-Arieh, Tom Zick, Kevin Klyman

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 24 pages, 3 figures, FAccT '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13466 2025-05-21 cs.AI 57%

AgentSGEN: Multi-Agent LLM in the Loop for Semantic Collaboration and GENeration of Synthetic Data

Vu Dinh Xuan, Hao Vo, David Murphy, Hoang D. Nguyen

机构 * University of Information Technology, VNU–HCM(越南胡志明市信息技术大学) University College Cork(科尔克大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12718 2025-05-20 cs.CL cs.HC 57%

Automated Bias Assessment in AI-Generated Educational Content Using CEAT Framework

Jingyang Peng, Wenyuan Shen, Jiarui Rao, Jionghao Lin

机构 * Carnegie Mellon University(卡内基梅隆大学) The University of Hong Kong(香港大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted by AIED 2025: Late-Breaking Results (LBR) Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12001 2025-05-20 cs.AI cs.MA 57%

Interactional Fairness in LLM Multi-Agent Systems: An Evaluation Framework

Ruta Binkyte

机构 * Ruta Binkyte(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07700 2025-05-15 cs.CY 57%

In Oxford Handbook on AI Governance: The Role of Workers in AI Ethics and Governance

Natalia Luka, JS Tan

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments In: Justin Bullock, Baobao Zhang, Yu-Che Chen, Johannes Himmelreich, Matthew Young, Antonin Korinek & Valerie Hudson (eds.). Oxford Handbook on AI Governance (Oxford University Press, 2022 forthcoming)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08404 2025-05-14 cs.AI 57%

Explaining Autonomous Vehicles with Intention-aware Policy Graphs

Sara Montese, Victor Gimenez-Abalos, Atia Cortés, Ulises Cortés, Sergio Alvarez-Napagao

机构 * Barcelona Supercomputing Center(巴塞罗那超级计算中心) Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Accepted to Workshop EXTRAAMAS 2025 in AAMAS Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07693 2025-05-13 cs.AI 57%

Belief Injection for Epistemic Control in Linguistic State Space

Sebastian Dumbrava

机构 * Sebastian Dumbrava(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 30 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07468 2025-05-13 cs.CY 57%

Promising Topics for U.S.-China Dialogues on AI Risks and Governance

Saad Siddiqui, Lujain Ibrahim, Kristy Loke, Stephen Clare, Marianne Lu, Aris Richardson, Conor McGlynn, Jeffrey Ding

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07005 2025-05-13 cs.AI 57%

Explainable AI the Latest Advancements and New Trends

Bowen Long, Enjie Liu, Renxi Qiu, Yanqing Duan

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04393 2025-05-08 cs.CL 57%

Large Means Left: Political Bias in Large Language Models Increases with Their Number of Parameters

David Exler, Mark Schutera, Markus Reischl, Luca Rettenberger

机构 * Institute for Automation and Applied Informatics(自动化与应用信息研究所) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04291 2025-05-08 cs.CY 57%

From Incidents to Insights: Patterns of Responsibility following AI Harms

Isabel Richards, Claire Benn, Miri Zilka

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03992 2025-05-08 cs.LG 57%

Algorithmic Accountability in Small Data: Sample-Size-Induced Bias Within Classification Metrics

Jarren Briscoe, Garrett Kepler, Daryl Deford, Assefaw Gebremedhin

机构 * School of Electrical Engineering & Computer Science, Washington State University(电气工程与计算机科学学院,华盛顿州立大学) Department of Mathematics and Statistics, Washington State University(数学与统计学系,华盛顿州立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

Comments AISTATS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02174 2025-05-06 cs.CY 57%

AI Governance in the GCC States: A Comparative Analysis of National AI Strategies

Mohammad Rashed Albous, Odeh Rashed Al-Jayyousi, Melodena Stephens

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments 33 pages,6 figures, 11 tables

Journal ref Journal of Artificial Intelligence Research, 82, 2389-2422 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00841 2025-05-05 cs.CR cs.AI 57%

From Texts to Shields: Convergence of Large Language Models and Cybersecurity

Tao Li, Ya-Ting Yang, Yunian Pan, Quanyan Zhu

机构 * New York University(纽约大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21400 2025-05-01 econ.GN cs.CL q-fin.EC 57%

Who Gets the Callback? Generative AI and Gender Bias

Sugat Chaturvedi, Rochana Chaturvedi

机构 * Ahmedabad University(阿赫迈达巴大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20215 2025-04-30 cs.CY cs.HC econ.GN q-fin.EC 57%

Exploring AI-powered Digital Innovations from A Transnational Governance Perspective: Implications for Market Acceptance and Digital Accountability Accountability

Claire Li, David Peter Wallis Freeborn

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Journal ref Proceedings of the UK Academy for Information Systems Conference 2025, Newcastle, UK. UKAIS

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08260 2025-04-18 cs.CL 57%

Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare

Yonchanok Khaokaew, Flora D. Salim, Andreas Züfle, Hao Xue, Taylor Anderson, C. Raina MacIntyre, Matthew Scotch, David J Heslop

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19363 2025-04-11 cs.CL 57%

Expressivity and Speech Synthesis

Andreas Triantafyllopoulos, Björn W. Schuller

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Published in Oxford Handbook of Expressivity in Language (in press)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05610 2025-04-09 cs.LG 57%

Fairness in Machine Learning-based Hand Load Estimation: A Case Study on Load Carriage Tasks

Arafat Rahman, Sol Lim, Seokhyun Chung

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02874 2025-04-07 cs.CL 57%

TheBlueScrubs-v1, a comprehensive curated medical dataset derived from the internet

Luis Felipe, Carlos Garcia, Issam El Naqa, Monique Shotande, Aakash Tripathi, Vivek Rudrapatna, Ghulam Rasool, Danielle Bitterman, Gilmer Valdes

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

Comments 22 pages, 8 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23056 2025-04-07 cs.CY 57%

Achieving Socio-Economic Parity through the Lens of EU AI Act

Arjun Roy, Stavroula Rizou, Symeon Papadopoulos, Eirini Ntoutsi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02074 2025-04-04 cs.HC cs.AI 57%

Trapped by Expectations: Functional Fixedness in LLM-Enabled Chat Search

Jiqun Liu, Jamshed Karimnazarov, Ryen W. White

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23566 2025-04-01 cs.CL 57%

When LLM Therapists Become Salespeople: Evaluating Large Language Models for Ethical Motivational Interviewing

Haein Kong, Seonghyeon Moon

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20414 2025-03-27 cs.LG 57%

Active Data Sampling and Generation for Bias Remediation

Antonio Maratea, Rita Perna

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09630 2025-03-25 cs.CY cs.HC 57%

What does AI consider praiseworthy?

Andrew J. Peterson

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments Forthcoming in AI and Ethics

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16436 2025-03-24 cs.HC cs.AI cs.RO 57%

Enhancing Human-Robot Collaboration through Existing Guidelines: A Case Study Approach

Yutaka Matsubara, Akihisa Morikawa, Daichi Mizuguchi, Kiyoshi Fujiwara

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09488 2025-03-21 eess.SY cs.LG cs.SY 57%

Intelligent Agricultural Greenhouse Control System Based on Internet of Things and Machine Learning

Cangqing Wang, Jiangchuan Gong

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12367 2025-03-18 cs.LG physics.ao-ph 57%

Integrating mobile and fixed monitoring data for high-resolution PM2.5 mapping using machine learning

Rui Xu, Dawen Yao, Yuzhuang Pian, Ruhui Cao, Yixin Fu, Xinru Yang, Ting Gan, Yonghong Liu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07496 2025-03-14 cs.CY 57%

Securing External Deeper-than-black-box GPAI Evaluations

Alejandro Tlaie, Jimmy Farrell

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09858 2025-03-14 cs.AI cs.GT cs.MA nlin.CD 57%

Media and responsible AI governance: a game-theoretic and LLM analysis

Nataliya Balabanova, Adeela Bashir, Paolo Bova, Alessio Buscemi, Theodor Cimpeanu, Henrique Correia da Fonseca, Alessandro Di Stefano, Manh Hong Duong, Elias Fernandez Domingos, Antonio Fernandes, The Anh Han, Marcus Krellner, Ndidi Bianca Ogbo, Simon T. Powers, Daniele Proverbio, Fernando P. Santos, Zia Ush Shamszaman, Zhao Song

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏