arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2502.02221 2025-06-12 cs.LG cs.AI stat.ML 62%

Bias Detection via Maximum Subgroup Discrepancy

Jiří Němeček, Mark Kozdoba, Illia Kryvoviaz, Tomáš Pevný, Jakub Mareček

机构 * Czech Technical University in Prague, Faculty of Electrical Engineering(捷克技术大学布拉格电子工程学院)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 6 figures

Journal ref Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08231 2025-06-11 cs.LG cs.AI cs.PF 62%

Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework

Melissa Estevez, Nisha Singh, Lauren Dyson, Blythe Adamson, Qianyu Yuan, Megan W. Hildner, Erin Fidyk, Olive Mbah, Farhad Khan, Kathi Seidl-Rathkopf, Aaron B. Cohen

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 3 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05728 2025-06-04 cs.CY cs.AI 62%

Political Neutrality in AI Is Impossible- But Here Is How to Approximate It

Jillian Fisher, Ruth E. Appel, Chan Young Park, Yujin Potter, Liwei Jiang, Taylor Sorensen, Shangbin Feng, Yulia Tsvetkov, Margaret E. Roberts, Jennifer Pan, Dawn Song, Yejin Choi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments Code: https://github.com/jfisher52/Approximation_Political_Neutrality

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19331 2025-06-02 cs.LG cs.CY 62%

Friends in Unexpected Places: Enhancing Local Fairness in Federated Learning through Clustering

Yifan Yang, Ali Payani, Parinaz Naghizadeh

机构 * The Ohio State University(俄亥俄州立大学) Cisco Research(思科研究) UC, San Diego(圣地亚哥大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21570 2025-05-29 cs.CY cs.AI 62%

Beyond Explainability: The Case for AI Validation

Dalit Ken-Dror Feldman, Daniel Benoliel

机构 * University of Haifa(海法大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21562 2025-05-29 cs.CY cs.AI cs.HC 62%

Enhancing Selection of Climate Tech Startups with AI -- A Case Study on Integrating Human and AI Evaluations in the ClimaTech Great Global Innovation Challenge

Jennifer Turliuk, Alejandro Sevilla, Daniela Gorza, Tod Hynes

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21537 2025-05-29 cs.CY cs.AI 62%

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models

Hao Sun, Yunyi Shen, Mihaela van der Schaar

机构 * University of Cambridge(剑桥大学) MIT(麻省理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20692 2025-05-28 cs.HC cs.AI cs.CL 62%

Can we Debias Social Stereotypes in AI-Generated Images? Examining Text-to-Image Outputs and User Perceptions

Saharsh Barve, Andy Mao, Jiayue Melissa Shi, Prerna Juneja, Koustuv Saha

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15481 2025-05-27 cs.CL cs.CY 62%

Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency

Yiran Liu, Ke Yang, Zehan Qi, Xiao Liu, Yang Yu, ChengXiang Zhai

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19776 2025-05-27 cs.CL cs.AI 62%

Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification

Akram Elbouanani, Evan Dufraisse, Adrian Popescu

机构 * Université Paris-Saclay, CEA, List(巴黎-萨克雷大学、欧洲原子能机构、List)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments To be published in the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18494 2025-05-27 cs.LG cs.AI 62%

FedHL: Federated Learning for Heterogeneous Low-Rank Adaptation via Unbiased Aggregation

Zihao Peng, Jiandian Zeng, Boyuan Li, Guo Li, Shengbo Chen, Tian Wang

机构 * Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai(人工智能与未来网络研究院,北京师范大学,珠海) School of Computer Science and Artificial Intelligence, Zhengzhou University(计算机科学与人工智能学院,郑州大学) School of Software, Nanchang University(软件学院,南昌大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06790 2025-05-23 cs.CY cs.CL cs.HC 62%

Large-scale moral machine experiment on large language models

Muhammad Shahrul Zaim bin Ahmad, Kazuhiro Takemoto

机构 * Department of Bioscience and Bioinformatics, Kyushu Institute of Technology(九州工科大学生物科学与生物信息学系) Faculty of Engineering and Technology, Multimedia University(多媒体大学工程与技术学院) Data Science and AI Research Center, Kyushu Institute of Technology(九州工科大学数据科学与人工智能研究中心)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

Comments 21 pages, 6 figures

Journal ref PLoS One 20, e0322776 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15031 2025-05-22 cs.CL cs.AI cs.HC cs.IR 62%

Are the confidence scores of reviewers consistent with the review content? Evidence from top conference proceedings in AI

Wenqing Wu, Haixu Xi, Chengzhi Zhang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref Scientometrics, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13973 2025-05-21 cs.CL cs.AI cs.CV 62%

Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models

Wenhui Zhu, Xuanzhao Dong, Xin Li, Peijie Qiu, Xiwen Chen, Abolfazl Razi, Aris Sotiras, Yi Su, Yalin Wang

机构 * Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学) Washington University in St.Louis(华盛顿大学圣路易斯分校) Banner Alzheimer’s Institute(Banner阿尔茨海默病研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12625 2025-05-20 cs.CL cs.CR cs.LG 62%

R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model

Ali Naseh, Harsh Chaudhari, Jaechul Roh, Mingshi Wu, Alina Oprea, Amir Houmansadr

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Northeastern University(东北大学) GFW Report(GFW报告)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15508 2025-05-16 cs.CL cs.AI 62%

Compensate Quantization Errors+: Quantized Models Are Inquisitive Learners

Yifei Gao, Jie Ou, Lei Wang, Jun Cheng, Mengchu Zhou

机构 * Yifei Gao ∗ , Jie Ou ∗ , Lei Wang † † {\dagger} † , Jun Cheng, and Mengchu Zhou ∗(作者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Effecient Quantization Methods for LLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09576 2025-05-15 cs.CY cs.AI 62%

Ethics and Persuasion in Reinforcement Learning from Human Feedback: A Procedural Rhetorical Approach

Shannon Lodoen, Alexi Orchard

专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.AI、cs.CY

Comments 10 pages, 1 figure, Accepted version

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08106 2025-05-14 cs.CL cs.AI 62%

Are LLMs complicated ethical dilemma analyzers?

Jiashen, Du, Jesse Yao, Allen Liu, Zhekai Zhang

机构 * Department of Computer Science, University of California, Berkeley(加州大学伯克利分校计算机科学系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments CS194-280 Advanced LLM Agents project. Project page: https://github.com/ALT-JS/ethicaLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08064 2025-05-14 cs.HC cs.AI cs.CY 62%

Justified Evidence Collection for Argument-based AI Fairness Assurance

Alpay Sabuncuoglu, Christopher Burr, Carsten Maple

机构 * The Alan Turing Institute United Kingdom(阿尔法·图灵研究所(英国)) University of Warwick United Kingdom(沃里克大学(英国))

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments The paper is accepted for ACM Conference on Fairness, Accountability, and Transparency (ACM FAccT '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07875 2025-05-14 cs.CY cs.AI 62%

Getting Ready for the EU AI Act in Healthcare. A call for Sustainable AI Development and Deployment

John Brandt Brodersen, Ilaria Amelia Caggiano, Pedro Kringen, Vince Istvan Madai, Walter Osika, Giovanni Sartor, Ellen Svensson, Magnus Westerlund, Roberto V. Zicari

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 8 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06326 2025-05-13 cs.CY cs.AI 62%

Enterprise Architecture as a Dynamic Capability for Scalable and Sustainable Generative AI adoption: Bridging Innovation and Governance in Large Organisations

Alexander Ettinger

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 82 pages excluding appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14362 2025-05-12 cs.HC cs.AI cs.CY 62%

The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health Support

Inhwa Song, Sachin R. Pendse, Neha Kumar, Munmun De Choudhury

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments The first two authors contributed equally to this work; typos corrected and post-review revisions incorporated

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05197 2025-05-09 cs.AI cs.CY 62%

Societal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quilt

Joel Z. Leibo, Alexander Sasha Vezhnevets, William A. Cunningham, Sébastien Krier, Manfred Diaz, Simon Osindero

机构 * Google DeepMind(谷歌DeepMind) University of Toronto(多伦多大学) Mila - Québec AI Institute(魁北克AI研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15941 2025-05-06 cs.CL cs.AI 62%

FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity

Fanny Jourdan, Yannick Chevalier, Cécile Favre

机构 * IRT Saint Exupery(IRT圣埃克苏佩里) Université Lumière Lyon 2(里莫大学 Lyon 2) Université Claude Bernard Lyon 1(克劳德·贝尔纳大学 Lyon 1) ERIC(埃里克)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments FAccT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04785 2025-05-06 cs.CL cs.CY 62%

Mapping Trustworthiness in Large Language Models: A Bibliometric Analysis Bridging Theory to Practice

José Siqueira de Cerqueira, Kai-Kristian Kemell, Rebekah Rousi, Nannan Xi, Juho Hamari, Pekka Abrahamsson

机构 * Tampere University(塔尔基耶大学) University of Vaasa(瓦萨大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00339 2025-05-02 cs.CL cs.AI 62%

Enhancing AI-Driven Education: Integrating Cognitive Frameworks, Linguistic Feedback Analysis, and Ethical Considerations for Improved Content Generation

Antoun Yaacoub, Sansiri Tarnpradab, Phattara Khumprom, Zainab Assaghir, Lionel Prevost, Jérôme Da-Rugna

机构 * Learning, Data and Robotics (LDR) ESIEA Lab(学习、数据与机器人(LDR)ESIEA实验室) Department of Computer Engineering(计算机工程系) Graduate School of Management and Innovation(管理与创新研究生院) Faculty of Science(科学学院) Lebanese University(黎巴嫩大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments This article will be presented in IJCNN 2025 "AI Innovations for Education: Transforming Teaching and Learning through Cutting-Edge Technologies" workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09495 2025-05-01 cs.LG cs.AI 62%

FADE: Towards Fairness-aware Generation for Domain Generalization via Classifier-Guided Score-based Diffusion Models

Yujie Lin, Dong Li, Minglai Shao, Guihong Wan, Chen Zhao

机构 * School of New Media and Communication, Tianjin University(天津大学新媒体与传播学院) School of Informatics, Xiamen University(厦门大学信息学院) Department of Computer Science, Baylor University(贝勒大学计算机科学系) Departments of Biostatistics and Epidemiology, Harvard University(哈佛大学生物统计学与流行病学系)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19255 2025-04-29 cs.AI cs.CY 62%

The Convergent Ethics of AI? Analyzing Moral Foundation Priorities in Large Language Models with a Multi-Framework Approach

Chad Coleman, W. Russell Neuman, Ali Dasdan, Safinah Ali, Manan Shah

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 25 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17401 2025-04-29 cs.CY cs.AI cs.HC 62%

AIJIM: A Scalable Model for Real-Time AI in Environmental Journalism

Torsten Tiltack

机构 * Torsten Tiltack(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 22 pages, 10 figures, 5 tables. Keywords: Artificial Intelligence, Environmental Journalism, Real-Time Reporting, Vision Transformers, Image Recognition, Crowdsourced Validation, GPT-4, Automated News Generation, GIS Integration, Data Privacy Compliance, Explainable AI (XAI), AI Ethics, Sustainable Development

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04476 2025-04-28 cs.CY cs.AI 62%

The Moral Mind(s) of Large Language Models

Avner Seror

机构 * Aix Marseille Univ, CNRS, AMSE, Marseille, France(阿维尼翁-马赛大学,国家科学研究中心,AMSE,马赛,法国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏