arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1832 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1832 篇

2508.16642 2025-08-26 cs.CY 57%

AI as IA: The use and abuse of artificial intelligence (AI) for human enhancement through intellectual augmentation (IA)

Alexandre Erler, Vincent C. Müller

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Journal ref (2023) in Marcello Ienca and Fabrice Jotterand (eds.), The Routledge Handbook of the Ethics of Human Enhancement (London: Routledge), 187-99

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16013 2025-08-25 cs.CL 57%

Political Ideology Shifts in Large Language Models

Pietro Bernardelle, Stefano Civelli, Leon Fröhling, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini

机构 * The University of Queensland(昆士兰大学) University of Udine(乌迪内大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14415 2025-08-21 cs.AI 57%

The Agent Behavior: Model, Governance and Challenges in the AI Digital Age

Qiang Zhang, Pei Yan, Yijia Xu, Chuanpo Fu, Yong Fang, Yang Liu

机构 * School of Cyber Science and Engineering, Sichuan University, China(计算机科学与工程学院,四川大学) College of Computing and Data Science, Nanyang Technological University, Sinapore(计算与数据科学学院,南洋理工大学) College of Electronics and Information Engineering, Shenzhen University, China(电子与信息工程学院,深圳大学) Department of Computer Science and Technology, Tsinghua University, China(计算机科学与技术系,清华大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13743 2025-08-20 cs.CL 57%

Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA

Kaiwei Zhang, Qi Jia, Zijian Chen, Wei Sun, Xiangyang Zhu, Chunyi Li, Dandan Zhu, Guangtao Zhai

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12174 2025-08-19 cs.CY 57%

Urban AI Governance Must Embed Legal Reasonableness for Democratic and Sustainable Cities

Rashid Mushkani

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11824 2025-08-19 cs.SE cs.AI cs.CR cs.PF 57%

Rethinking Autonomy: Preventing Failures in AI-Driven Software Engineering

Satyam Kumar Navneet, Joydeep Chandra

机构 * Department of CSE Chandigarh University Mohali, India(计算机科学与工程系 奇纳格里大学 莫哈利,印度) Department of CST Tsinghua University Beijing, China(计算机科学与技术系 清华大学 北京,中国)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11262 2025-08-18 cs.CV cs.AI 57%

Vision-Language Models display a strong gender bias

Aiswarya Konavoor, Raj Abhijit Dandekar, Rajat Dandekar, Sreedath Panat

机构 * Togo AI Labs(Togo人工智能实验室) Vizuara AI Labs(Vizuara人工智能实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10007 2025-08-15 cs.CL stat.ME 57%

Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models

Y. Lyu, D. Combs, D. Neumann, Y. C. Leong

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments We have no known conflict of interest

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09019 2025-08-13 cs.AI 57%

Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs

Shivam Dubey

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08262 2025-08-13 cs.CL 57%

Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models

Alaa Alhamzeh, Mays Al Rebdawi

机构 * First Author Affiliation(第一作者机构) Second Author Affiliation(第二作者机构)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments 8 pages, 4 figures, Passau uni, Master thesis in NLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16170 2025-08-12 cs.AI 57%

Learning How to Vote with Principles: Axiomatic Insights Into the Collective Decisions of Neural Networks

Levin Hornischer, Zoi Terzopoulou

机构 * Munich Center for Mathematical Philosophy, LMU Munich Munich Germany GATE, CNRS, Universit\'e Jean Monnet, Universit\'e Lumiere Lyon 2 Saint-Etienne France Munich Center for Mathematical Philosophy, LMU Munich GATE, CNRS, Universit\'e Jean Monnet, Universit\'e Lumiere Lyon 2

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 44 pages, 21 figures, 14 tables. Updated and published version

Journal ref Journal of Artificial Intelligence Research 83, Article 25 (August 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06479 2025-08-11 cs.CY 57%

The Problem of Atypicality in LLM-Powered Psychiatry

Bosco Garcia, Eugene Y. S. Chua, Harman Singh Brah

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Preprint of 8/8/2025 -- please cite published version. This article has been published in the Journal of Medical Ethics (2025) following peer review and can also be viewed on the journal's website at 10.1136/jme-2025-110972

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04071 2025-08-07 cs.LG 57%

Adversarial Fair Multi-View Clustering

Mudi Jiang, Jiahui Zhou, Lianyu Hu, Xinying Liu, Zengyou He, Zhikui Chen

机构 * School of Software, Dalian University of Technology(大连理工大学软件学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17945 2025-08-07 cs.CL 57%

Assessing Agentic Large Language Models in Multilingual National Bias

Qianying Liu, Katrina Qiyao Wang, Fei Cheng, Sadao Kurohashi

机构 * National Institute of Informatics, Japan(日本信息机构国家研究所) University of Wisconsin—Madison, USA(美国威斯康星大学麦迪逊分校) Kyoto University, Japan(日本京都大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted to ACL 2025 Findings. 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10945 2025-08-07 cs.LG stat.ML 57%

Gradient-Based Multi-Objective Deep Learning: Algorithms, Theories, Applications, and Beyond

Weiyu Chen, Baijiong Lin, Xiaoyuan Zhang, Xi Lin, Han Zhao, Qingfu Zhang, James T. Kwok

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) City University of Hong Kong(香港城市大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00847 2025-08-05 cs.HC cs.CY 57%

GPT Chatbots for Alleviating Anxiety and Depression: A Pilot Randomized Controlled Trial with Afghan Women

Sofia Sahab, Jawad Haqbeen, Diksha Sapkota, Takayuki Ito

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23454 2025-08-04 cs.HC cs.CY cs.ET cs.GR q-bio.NC 57%

Breaking the mould of Social Mixed Reality - State-of-the-Art and Glossary

Marta Bieńkiewicz, Julia Ayache, Panayiotis Charalambous, Cristina Becchio, Marco Corragio, Bertram Taetz, Francesco De Lellis, Antonio Grotta, Anna Server, Daniel Rammer, Richard Kulpa, Franck Multon, Azucena Garcia-Palacios, Jessica Sutherland, Kathleen Bryson, Stéphane Donikian, Didier Stricker, Benoît Bardy

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21929 2025-07-30 cs.AI 57%

Libra: Large Chinese-based Safeguard for AI Content

Ziyang Chen, Huimu Yu, Xing Wu, Dongqin Liu, Songlin Hu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21319 2025-07-30 cs.CL 57%

Do Large Language Models Understand Morality Across Cultures?

Hadi Mohammadi, Yasmeen F. S. S. Meijer, Efthymia Papadopoulou, Ayoub Bagheri

机构 * Department of Methodology and Statistics, Utrecht University, The Netherlands(方法论与统计学系,乌特雷赫特大学,荷兰)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19962 2025-07-29 cs.CL 57%

KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models

Seorin Kim, Dongyoung Lee, Jaejin Lee

机构 * Dept. of Data Science, Seoul National University(数据科学系,首尔国立大学) Dept. of Computer Science and Engineering, Seoul National University(计算机科学与工程系,首尔国立大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13138 2025-07-29 cs.CL 57%

Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation

Hadi Mohammadi, Tina Shahedi, Pablo Mosteiro, Massimo Poesio, Ayoub Bagheri, Anastasia Giachanou

机构 * Department of Methodology and Statistics, Utrecht University, The Netherlands(方法论与统计学系,乌特列支大学,荷兰) Department of Information and Computing Sciences, Utrecht University, The Netherlands(信息与计算科学系,乌特列支大学,荷兰) Queen Mary University of London, London, United Kingdom(伦敦女王玛丽大学,伦敦,英国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13175 2025-07-29 cs.AI 57%

Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era

Matthew E. Brophy

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 42 pages. Supplementary material included at end of article

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17368 2025-07-24 cs.LG 57%

ViRN: Variational Inference and Distribution Trilateration for Long-Tailed Continual Representation Learning

Hao Dai, Chong Tang, Jagmohan Chauhan

机构 * Department of Computer Science, UCL Centre for Artificial Intelligence, University College London, London, UK(计算机科学系,UCL人工智能中心,伦敦大学学院,伦敦,英国) University of Southampton, Southampton, UK(南安普顿大学,南安普顿,英国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13294 2025-07-24 cs.CY cs.HC 57%

The "Who", "What", and "How" of Responsible AI Governance: A Systematic Review and Meta-Analysis of (Actor, Stage)-Specific Tools

Blaine Kuehnert, Rachel M. Kim, Jodi Forlizzi, Hoda Heidari

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments Accepted to ACM Conference on Fairness, Accountability, and Transparency 2025. 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14226 2025-07-23 cs.CY 57%

Mapping the Parasocial AI Market: User Trends, Engagement and Risks

Zilan Qian, Mari Izumikawa, Fiona Lodge, Angelo Leone

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 17 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14332 2025-07-22 cs.LG 57%

Development and Deployment of Hybrid ML Models for Critical Heat Flux Prediction in Annulus Geometries

Aidan Furlong, Xingang Zhao, Robert Salko, Xu Wu

机构 * Department of Nuclear Engineering, North Carolina State University(核工程系,北卡罗来纳州立大学) Department of Nuclear Engineering, University of Tennessee(核工程系,田纳西大学) Nuclear Energy and Fuel Cycle Division, Oak Ridge National Laboratory(核能与燃料循环 division,橡树岭国家实验室)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments Accepted for inclusion in Transactions of the American Nuclear Society for the 2025 ANS Winter Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13616 2025-07-21 cs.HC cs.CY cs.ET cs.IT cs.MA math.IT 57%

From Firms to Computation: AI Governance and the Evolution of Institutions

Michael S. Harre

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments 44 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19232 2025-07-18 cs.IR cs.AI 57%

LLM-RecG: A Semantic Bias-Aware Framework for Zero-Shot Sequential Recommendation

Yunzhe Li, Junting Wang, Hari Sundaram, Zhining Liu

机构 * University of Illinois, Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 10 pages, Recsys'25 Spotlight Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11344 2025-07-16 cs.LG 57%

Guiding LLM Decision-Making with Fairness Reward Models

Zara Hall, Melanie Subbiah, Thomas P Zollo, Kathleen McKeown, Richard Zemel

机构 * Columbia University(哥伦比亚大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10861 2025-07-16 cs.LG 57%

Visually grounded emotion regulation via diffusion models and user-driven reappraisal

Edoardo Pinzuti, Oliver Tüscher, André Ferreira Castro

机构 * Leibniz Institute for Resilience Research(莱比锡韧性研究所) University Medicine Halle (Saale) of the Martin Luther University Halle-Wittenberg (MLU)(马尔堡-哈雷大学哈雷-萨勒医学院) German Center for Mental Health (DZPG)(德国心理健康中心) School of Life Sciences, Technical University of Munich(慕尼黑技术大学生命科学学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏