arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1832 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1832 篇

2405.07076 2024-05-15 cs.CL cs.AI 62%

Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models

Edward Y. Chang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 29 pages, 10 tables, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14660 2024-04-24 cs.CY cs.AI 62%

AI Procurement Checklists: Revisiting Implementation in the Age of AI Governance

Tom Zick, Mason Kortz, David Eaves, Finale Doshi-Velez

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11271 2024-04-02 cs.CL cs.CY cs.HC 62%

MONAL: Model Autophagy Analysis for Modeling Human-AI Interactions

Shu Yang, Muhammad Asif Ali, Lu Yu, Lijie Hu, Di Wang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06153 2024-03-29 cs.LG cs.AI cs.HC 62%

Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments

Anthony Cintron Roman, Jennifer Wortman Vaughan, Valerie See, Steph Ballard, Jehu Torres, Caleb Robinson, Juan M. Lavista Ferres

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17368 2024-03-27 cs.CL cs.AI 62%

ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?

Fan Huang, Haewoon Kwak, Kunwoo Park, Jisun An

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accpeted by LREC-COLING 2024 main conference, long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15601 2024-03-26 cs.CY cs.AI 62%

From Guidelines to Governance: A Study of AI Policies in Education

Aashish Ghimire, John Edwards

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14469 2024-03-22 cs.CL cs.AI 62%

ChatGPT Alternative Solutions: Large Language Models Survey

Hanieh Alipour, Nick Pendar, Kohinoor Roy

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref David C. Wyld et al. (Eds): NBIoT, MLCL, NMCO, ARIN, CSITA, ISPR, NATAP-2024. pp. 153-173, 2024. CS & IT - CSCP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13840 2024-03-22 cs.CL cs.AI cs.SI 62%

Whose Side Are You On? Investigating the Political Stance of Large Language Models

Pagnarasmey Pit, Xingjun Ma, Mike Conway, Qingyu Chen, James Bailey, Henry Pit, Putrasmey Keo, Watey Diep, Yu-Gang Jiang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11402 2024-03-19 cs.CY cs.AI 62%

Embracing the Generative AI Revolution: Advancing Tertiary Education in Cybersecurity with GPT

Raza Nowrozy, David Jam

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09510 2024-03-15 cs.AI cs.CY cs.GT cs.MA math.DS 62%

Trust AI Regulation? Discerning users are vital to build trust and effective AI regulation

Zainab Alalawi, Paolo Bova, Theodor Cimpeanu, Alessandro Di Stefano, Manh Hong Duong, Elias Fernandez Domingos, The Anh Han, Marcus Krellner, Bianca Ogbo, Simon T. Powers, Filippo Zimmaro

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00064 2024-03-01 cs.CY cs.AI 62%

Ethical Framework for Harnessing the Power of AI in Healthcare and Beyond

Sidra Nasir, Rizwan Ahmed Khan, Samita Bai

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Journal ref IEEE Access 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13379 2024-02-22 cs.LG cs.CY 62%

Referee-Meta-Learning for Fast Adaptation of Locational Fairness

Weiye Chen, Yiqun Xie, Xiaowei Jia, Erhu He, Han Bao, Bang An, Xun Zhou

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04489 2024-02-08 cs.LG cs.CR cs.CY stat.ME 62%

De-amplifying Bias from Differential Privacy in Language Model Fine-tuning

Sanjari Srivastava, Piotr Mardziel, Zhikhun Zhang, Archana Ahlawat, Anupam Datta, John C Mitchell

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07353 2024-02-06 cs.SE cs.AI cs.LG 62%

Towards Engineering Fair and Equitable Software Systems for Managing Low-Altitude Airspace Authorizations

Usman Gohar, Michael C. Hunter, Agnieszka Marczak-Czajka, Robyn R. Lutz, Myra B. Cohen, Jane Cleland-Huang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Journal ref ICSE-SEIS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00402 2024-02-02 cs.CL cs.AI 62%

Investigating Bias Representations in Llama 2 Chat via Activation Steering

Dawn Lu, Nina Rimsky

专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16088 2024-01-30 cs.LG cs.CY 62%

Fairness in Algorithmic Recourse Through the Lens of Substantive Equality of Opportunity

Andrew Bell, Joao Fonseca, Carlo Abrate, Francesco Bonchi, Julia Stoyanovich

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10310 2024-01-22 cs.LG cs.AI cs.CC 62%

Mathematical Algorithm Design for Deep Learning under Societal and Judicial Constraints: The Algorithmic Transparency Requirement

Holger Boche, Adalbert Fono, Gitta Kutyniok

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09473 2024-01-19 cs.CY cs.AI 62%

Business and ethical concerns in domestic Conversational Generative AI-empowered multi-robot systems

Rebekah Rousi, Hooman Samani, Niko Mäkitalo, Ville Vakkuri, Simo Linkola, Kai-Kristian Kemell, Paulius Daubaris, Ilenia Fronza, Tommi Mikkonen, Pekka Abrahamsson

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 15 pages, 4 figures, International Conference on Software Business

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06709 2024-01-15 cs.CL cs.AI 62%

Reliability Analysis of Psychological Concept Extraction and Classification in User-penned Text

Muskan Garg, MSVPJ Sathvik, Amrit Chadha, Shaina Raza, Sunghwan Sohn

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18333 2023-12-18 cs.CL cs.AI 62%

She had Cobalt Blue Eyes: Prompt Testing to Create Aligned and Sustainable Language Models

Veronica Chatrath, Oluwanifemi Bamgbose, Shaina Raza

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted as Oral at the AAAI 2nd Workshop on Sustainable AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15936 2023-12-01 cs.CY cs.LG 62%

Towards Responsible Governance of Biological Design Tools

Richard Moulange, Max Langenkamp, Tessa Alexanian, Samuel Curtis, Morgan Livingston

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

Comments 10 pages + references, 1 figure, accepted at NeurIPS 2023 Workshop on Regulatable ML as oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14684 2023-11-28 cs.CY cs.AI 62%

The risks of risk-based AI regulation: taking liability seriously

Martin Kretschmer, Tobias Kretschmer, Alexander Peukert, Christian Peukert

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07141 2023-11-17 cs.LG cs.CY 62%

SABAF: Removing Strong Attribute Bias from Neural Networks with Adversarial Filtering

Jiazhi Li, Mahyar Khayatkhoei, Jiageng Zhu, Hanchen Xie, Mohamed E. Hussein, Wael AbdAlmageed

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

Comments 35 pages, 18 figures, 32 tables. This work is an extended version of our paper (arXiv:2310.04955). Code will be released at https://github.com/jiazhi412/strong_attribute_bias

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02294 2023-11-07 cs.CL cs.CY 62%

LLMs grasp morality in concept

Mark Pock, Andre Ye, Jared Moore

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

Comments Presented at NeurIPS 2023 Moral Pyschology and Moral Philosophy workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08901 2023-10-16 cs.MA cs.AI cs.CL 62%

Welfare Diplomacy: Benchmarking Language Model Cooperation

Gabriel Mukobi, Hannah Erlebach, Niklas Lauffer, Lewis Hammond, Alan Chan, Jesse Clifton

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02796 2023-10-12 cs.DB cs.CL cs.LG 62%

VerifAI: Verified Generative AI

Nan Tang, Chenyu Yang, Ju Fan, Lei Cao, Yuyu Luo, Alon Halevy

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.LG

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05840 2023-10-10 cs.LG cs.AI 62%

Predicting Accident Severity: An Analysis Of Factors Affecting Accident Severity Using Random Forest Model

Adekunle Adefabi, Somtobe Olisah, Callistus Obunadike, Oluwatosin Oyetubo, Esther Taiwo, Edward Tella

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments 15 pages

Journal ref International Journal on Cybernetics & Informatics (IJCI) Vol.12, No.6, December 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.04963 2023-09-29 cs.AI cs.CY cs.SE 62%

Responsible AI Pattern Catalogue: A Collection of Best Practices for AI Governance and Engineering

Qinghua Lu, Liming Zhu, Xiwei Xu, Jon Whittle, Didar Zowghi, Aurelie Jacquet

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12356 2023-09-25 cs.CY cs.AI 62%

A Critical Examination of the Ethics of AI-Mediated Peer Review

Laurie A. Schintler, Connie L. McNeely, James Witte

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 21 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10318 2023-09-20 cs.AI cs.CY 62%

Who to Trust, How and Why: Untangling AI Ethics Principles, Trustworthiness and Trust

Andreas Duenser, David M. Douglas

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 7 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏