arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1832 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1832 篇

2507.08979 2025-07-15 cs.CV cs.LG 57%

PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection

Mahdiyar Molahasani, Azadeh Motamedi, Michael Greenspan, Il-Min Kim, Ali Etemad

机构 * Queen’s University(皇后大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08594 2025-07-14 cs.SE cs.AI cs.HC 57%

Generating Proto-Personas through Prompt Engineering: A Case Study on Efficiency, Effectiveness and Empathy

Fernando Ayach, Vitor Lameirão, Raul Leão, Jerfferson Felizardo, Rafael Sobrinho, Vanessa Borges, Patrícia Matsubara, Awdren Fontão

机构 * Faculty of Computing - Federal University of Mato Grosso do Sul(计算机学院 - 莫扎尔河大省联邦大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 12 pages; 2 figures; Preprint with the original submission accepted for publication at 39th Brazilian Symposium on Software Engineering (SBES)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07604 2025-07-11 cs.LG q-bio.QM q-bio.TO 57%

Synthetic MC via Biological Transmitters: Therapeutic Modulation of the Gut-Brain Axis

Sebastian Lotter, Elisabeth Mohr, Andrina Rutsch, Lukas Brand, Francesca Ronchi, Laura Díaz-Marugán

机构 * Charité – Universitätsmedizin Berlin, Humboldt-Universität zu Berlin, Berlin Institute of Health (BIH), Berlin, Germany(柏林查理医院、洪堡-柏林大学、柏林健康研究所(BIH)、柏林)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05866 2025-07-09 cs.CY 57%

Understanding support for AI regulation: A Bayesian network perspective

Andrea Cremaschi, Dae-Jin Lee, Manuele Leonelli

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03458 2025-07-08 cs.CV cs.AI 57%

Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach

Leyan Xue, Zongbo Han, Guangyu Wang, Qinghua Hu, Mingyue Cheng, Changqing Zhang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22372 2025-06-30 cs.IR cs.CL 57%

Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement

Maryam Mousavian, Zahra Abbasiantaeb, Mohammad Aliannejadi, Fabio Crestani

机构 * Università della Svizzera italiana \& University of Amsterdam University of Amsterdam The Netherland Unviersity of Amsterdam The Netherland University of Amsterdam

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted by ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20743 2025-06-27 cs.LG cs.CE 57%

A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools

Minh-Hao Van, Prateek Verma, Chen Zhao, Xintao Wu

机构 * Department of EECS University of Arkansas Fayetteville(电子工程与计算机科学系美国阿肯色大学弗莱维尔分校) Department of CS Baylor University Waco(计算机科学系贝勒大学沃斯堡)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18732 2025-06-24 cs.LG 57%

Towards Group Fairness with Multiple Sensitive Attributes in Federated Foundation Models

Yuning Yang, Han Yu, Tianrun Gao, Xiaodong Xu, Guangyu Wang

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18802 2025-06-24 cs.CL 57%

Language Models Grow Less Humanlike beyond Phase Transition

Tatsuya Aoyama, Ethan Wilcox

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted to ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04345 2025-06-23 cs.CY 57%

Build Agent Advocates, Not Platform Agents

Sayash Kapoor, Noam Kolt, Seth Lazar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Accepted to ICML 2025 position paper track

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04686 2025-06-19 cs.AI 57%

Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization

Zelai Xu, Wanjun Gu, Chao Yu, Yi Wu, Yu Wang

机构 * Tsinghua University, Beijing, China(清华大学) Beijing Zhongguancun Academy, Beijing, China(北京中关村学院) Shanghai Qi Zhi Institute, Shanghai, China(上海启智研究所)

专题命中 AI治理与伦理 :DPO(abstract);分类 cs.AI

Comments Published in ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12527 2025-06-17 cs.CL 57%

Detection, Classification, and Mitigation of Gender Bias in Large Language Models

Xiaoqing Cheng, Hongying Zan, Lulu Kong, Jinwang Song, Min Peng

机构 * Zhengzhou University(郑州大学) Wuhan University(武汉大学)

专题命中 AI治理与伦理 :DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12245 2025-06-17 cs.AI 57%

Reversing the Paradigm: Building AI-First Systems with Human Guidance

Cosimo Spera, Garima Agrawal

机构 * Minerva CQ

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11068 2025-06-16 cs.CL 57%

Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models

Bumjin Park, Jinsil Lee, Jaesik Choi

机构 * KAIST AI(韩国科学技术院人工智能研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments 20 pages including references and appendix; To appear in ACL 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06366 2025-06-13 q-bio.NC cs.CY cs.MA 57%

AI Agent Behavioral Science

Lin Chen, Yunke Zhang, Jie Feng, Haoye Chai, Honglin Zhang, Bingbing Fan, Yibo Ma, Shiyuan Zhang, Nian Li, Tianhui Liu, Nicholas Sukiennik, Keyu Zhao, Yu Li, Ziyi Liu, Fengli Xu, Yong Li

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02348 2025-06-13 cs.LG stat.ML 57%

Simplicity bias and optimization threshold in two-layer ReLU networks

Etienne Boursier, Nicolas Flammarion

机构 * INRIA, LMO, Université Paris-Saclay, Orsay, France(INRIA、LMO、巴黎-萨克雷大学、欧萨斯分校、法国) TML Lab, EPFL, Switzerland(TML实验室、瑞士联邦理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments ICML camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08593 2025-06-11 cs.CL 57%

Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models

Shuzhou Yuan, Ercong Nie, Mario Tawfelis, Helmut Schmid, Hinrich Schütze, Michael Färber

机构 * ScaDS.AI and TU Dresden(ScaDS.AI 和 梵高大学) LMU Munich(慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06225 2025-06-10 cs.HC cs.AI 57%

"We need to avail ourselves of GenAI to enhance knowledge distribution": Empowering Older Adults through GenAI Literacy

Eunhye Grace Ko, Shaini Nanayakkara, Earl W. Huff

机构 * School of Information University of Texas at Austin(信息学院得克萨斯大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Journal ref CHI EA ' 2025: Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02923 2025-06-05 cs.AI stat.ML 57%

The Limits of Predicting Agents from Behaviour

Alexis Bellot, Jonathan Richens, Tom Everitt

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Journal ref ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02295 2025-06-04 cs.CL 57%

Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection

Yachao Zhao, Bo Wang, Yan Wang, Dongming Zhao, Ruifang He, Yuexian Hou

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) AI Lab, China Mobile Communication Group Tianjin Co., Ltd.(中国移动通信集团天津有限公司人工智能实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00616 2025-06-03 cs.CY 57%

Catastrophic Liability: Managing Systemic Risks in Frontier AI Development

Aidan Kierans, Kaley Rittichier, Utku Sonsayar, Avijit Ghosh

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 10 pages, 1 figure, 1 table, in review for AIES 2025, presented at TAIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21997 2025-05-29 cs.CL 57%

Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data

Jihong Zhang, Xinya Liang, Anqi Deng, Nicole Bonge, Lin Tan, Ling Zhang, Nicole Zarrett

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05787 2025-05-29 cs.CY 57%

Mapping the Regulatory Learning Space for the EU AI Act

Dave Lewis, Marta Lasek-Markey, Delaram Golpayegani, Harshvardhan J. Pandit

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19915 2025-05-28 cs.CR cs.AI 57%

Evaluating AI cyber capabilities with crowdsourced elicitation

Artem Petrov, Dmitrii Volkov

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Updated abstract to fix a typo; no changes to the content of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20273 2025-05-27 cs.AI 57%

Ten Principles of AI Agent Economics

Ke Yang, ChengXiang Zhai

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18657 2025-05-27 cs.AI 57%

MLLMs are Deeply Affected by Modality Bias

Xu Zheng, Chenfei Liao, Yuqian Fu, Kaiyu Lei, Yuanhuiyi Lyu, Lutao Jiang, Bin Ren, Jialei Chen, Jiawen Wang, Chengxin Li, Linfeng Zhang, Danda Pani Paudel, Xuanjing Huang, Yu-Gang Jiang, Nicu Sebe, Dacheng Tao, Luc Van Gool, Xuming Hu

机构 * HKUST(GZ)(香港科技大学(广州)) CSE, HKUST(香港科技大学计算机科学与工程系) Xi’an Jiaotong University(西安交通大学) University of Pisa, IT(比萨大学) University of Trento, IT(特伦特大学) Nagoya University(名古屋大学) China University of Mining & Technology, Beijing(中国矿业大学(北京)) Tongji University(同济大学) SPIC Energy Science and Technology Research Institute(SPIC能源科学与技术研究院) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18523 2025-05-27 cs.CY 57%

Diversity and Inclusion in AI: Insights from a Survey of AI/ML Practitioners

Sidra Malik, Muneera Bano, Didar Zowghi

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13981 2025-05-27 cs.CV cs.AI 57%

On the Fairness, Diversity and Reliability of Text-to-Image Generative Models

Jordan Vice, Naveed Akhtar, Leonid Sigal, Richard Hartley, Ajmal Mian

机构 * University of Western Australia(西澳大学) University of Melbourne(墨尔本大学) University of British Columbia(不列颠哥伦比亚大学) Australian National University(澳大利亚国立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments This research is supported by the NISDRG project #20100007, funded by the Australian Government

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17712 2025-05-26 cs.CL 57%

Understanding How Value Neurons Shape the Generation of Specified Values in LLMs

Yi Su, Jiayi Zhang, Shu Yang, Xinhai Wang, Lijie Hu, Di Wang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19294 2025-05-23 cs.LG 57%

Investigating the Effects of Fairness Interventions Using Pointwise Representational Similarity

Camila Kolling, Till Speicher, Vedant Nanda, Mariya Toneva, Krishna P. Gummadi

机构 * MPI-SWS(马克斯·普朗克所社会科学研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏