arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1832 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1832 篇

2410.13250 2024-10-18 cs.HC cs.AI cs.CY 62%

Perceptions of Discriminatory Decisions of Artificial Intelligence: Unpacking the Role of Individual Characteristics

Soojong Kim

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13138 2024-10-18 cs.CL cs.CR cs.CY 62%

Data Defenses Against Large Language Models

William Agnew, Harry H. Jiang, Cella Sum, Maarten Sap, Sauvik Das

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12848 2024-10-18 cs.CL cs.AI cs.HC 62%

Prompt Engineering a Schizophrenia Chatbot: Utilizing a Multi-Agent Approach for Enhanced Compliance with Prompt Instructions

Per Niklas Waaler, Musarrat Hussain, Igor Molchanov, Lars Ailo Bongo, Brita Elvevåg

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08760 2024-10-16 cs.CL cs.AI 62%

The Generation Gap: Exploring Age Bias in the Value Systems of Large Language Models

Siyang Liu, Trish Maturi, Bowen Yi, Siqi Shen, Rada Mihalcea

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 5 pages

Journal ref The 2024 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12232 2024-10-08 cs.AI cs.CL 62%

"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations

Huy Nghiem, John Prindle, Jieyu Zhao, Hal Daumé

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024, 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01819 2024-10-04 cs.CY cs.AI 62%

Strategic AI Governance: Insights from Leading Nations

Dian W. Tjondronegoro

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 21 pages, 3 Figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12926 2024-09-24 cs.LG cs.AI 62%

Trusting Fair Data: Leveraging Quality in Fairness-Driven Data Removal Techniques

Manh Khoi Duong, Stefan Conrad

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments The Version of Record of this contribution is published in Springer LNCS 14912 and is available online at https://doi.org/10.1007/978-3-031-68323-7_33

Journal ref Lecture Notes in Computer Science, Vol. 14912 (2024), pp. 375-380. Springer

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13852 2024-09-24 cs.CL cs.AI 62%

Do language models practice what they preach? Examining language ideologies about gendered language reform encoded in LLMs

Julia Watson, Sophia Lee, Barend Beekhuizen, Suzanne Stevenson

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07335 2024-09-12 cs.AI cs.CL 62%

Explanation, Debate, Align: A Weak-to-Strong Framework for Language Model Generalization

Mehrdad Zakershahrak, Samira Ghodratnama

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07467 2024-09-09 cs.CY cs.AI 62%

AI Ethics Principles in Practice: Perspectives of Designers and Developers

Conrad Sanderson, David Douglas, Qinghua Lu, Emma Schleiger, Jon Whittle, Justine Lacey, Glenn Newnham, Stefan Hajkowicz, Cathy Robinson, David Hansen

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments submitted to IEEE Transactions on Technology & Society

Journal ref IEEE Transactions on Technology and Society, Vol. 4, No. 2, pp. 171-187, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12691 2024-09-04 cs.AI cs.CY 62%

Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?

Shayne Longpre, Robert Mahari, Naana Obeng-Marnu, William Brannon, Tobin South, Katy Gero, Sandy Pentland, Jad Kabbara

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments ICML 2024 camera-ready version (Spotlight paper). 9 pages, 2 tables

Journal ref Proceedings of ICML 2024, in PMLR 235:32711-32725. URL: https://proceedings.mlr.press/v235/longpre24b.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12055 2024-08-23 cs.CL cs.LG 62%

Aligning (Medical) LLMs for (Counterfactual) Fairness

Raphael Poulain, Hamed Fayyaz, Rahmatollah Beheshti

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:2404.15149

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11457 2024-08-22 cs.IR cs.AI cs.CL 62%

Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era

Sunhao Dai, Chen Xu, Shicheng Xu, Liang Pang, Zhenhua Dong, Jun Xu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments KDD 2024 Tutorial&Survey; Tutorial Website: https://llm-ir-bias-fairness.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10608 2024-08-21 cs.CL cs.AI 62%

Promoting Equality in Large Language Models: Identifying and Mitigating the Implicit Bias based on Bayesian Theory

Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Jing Pan, Chen Jue, Zhijun Fang, Yinghui Xu, Wei Chu, Yuan Qi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03907 2024-08-08 cs.CL cs.AI 62%

Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models

Shachi H Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Beckage, Hsuan Su, Hung-yi Lee, Lama Nachman

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 6 pages paper content, 17 pages of appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21281 2024-08-01 cs.CY cs.AI 62%

Unlocking the Potential of Binding Corporate Rules (BCRs) in Health Data Transfers

Marcelo Corrales Compagnucci, Mark Fenwick, Helena Haapio

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16903 2024-07-25 cs.CY cs.AI 62%

US-China perspectives on extreme AI risks and global governance

Akash Wasil, Tim Durgin

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13025 2024-07-22 cs.CY cs.AI 62%

From Principles to Practices: Lessons Learned from Applying Partnership on AI's (PAI) Synthetic Media Framework to 11 Use Cases

Claire R. Leibowicz, Christian H. Cardona

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 18 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12999 2024-07-19 cs.CY cs.AI cs.CR 62%

Securing the Future of GenAI: Policy and Technology

Mihai Christodorescu, Ryan Craven, Soheil Feizi, Neil Gong, Mia Hoffmann, Somesh Jha, Zhengyuan Jiang, Mehrdad Saberi Kamarposhti, John Mitchell, Jessica Newman, Emelia Probasco, Yanjun Qi, Khawaja Shams, Matthew Turek

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12273 2024-07-11 cs.CR cs.AI cs.CL 62%

The Ethics of Interaction: Mitigating Security Threats in LLMs

Ashutosh Kumar, Shiv Vignesh Murthy, Sagarika Singh, Swathy Ragupathy

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06235 2024-07-10 cs.CY cs.AI 62%

Auditing of AI: Legal, Ethical and Technical Approaches

Jakob Mokander

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Journal ref DISO 2, 49 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19497 2024-07-01 cs.CL cs.AI 62%

Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts

Naseela Pervez, Alexander J. Titus

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09779 2024-06-17 cs.AI cs.CL cs.CV 62%

OSPC: Detecting Harmful Memes with Large Language Model as a Catalyst

Jingtao Cao, Zheng Zhang, Hongru Wang, Bin Liang, Hao Wang, Kam-Fai Wong

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06435 2024-06-11 cs.CL cs.AI 62%

Language Models are Alignable Decision-Makers: Dataset and Application to the Medical Triage Domain

Brian Hu, Bill Ray, Alice Leung, Amy Summerville, David Joy, Christopher Funk, Arslan Basharat

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 15 pages total (including appendix), NAACL 2024 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10819 2024-06-11 cs.LG cs.AI stat.ML 62%

Auditing and Generating Synthetic Data with Controllable Trust Trade-offs

Brian Belgodere, Pierre Dognin, Adam Ivankay, Igor Melnyk, Youssef Mroueh, Aleksandra Mojsilovic, Jiri Navratil, Apoorva Nitsure, Inkit Padhi, Mattia Rigotti, Jerret Ross, Yair Schiff, Radhika Vedpathak, Richard A. Young

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04671 2024-06-10 cs.CY cs.AI 62%

The Reasonable Person Standard for AI

Sunayana Rane

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03299 2024-06-06 cs.AI cs.CL 62%

The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional Decisions of Large Language Models in Cooperation and Bargaining Games

Mikhail Mozikov, Nikita Severin, Valeria Bodishtianu, Maria Glushanina, Mikhail Baklashkin, Andrey V. Savchenko, Ilya Makarov

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13041 2024-06-06 cs.CL cs.AI 62%

Assessing Political Bias in Large Language Models

Luca Rettenberger, Markus Reischl, Mark Schutera

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01264 2024-05-24 cs.AI cs.CL 62%

Exploring the psychology of LLMs' Moral and Legal Reasoning

Guilherme F. C. F. Almeida, José Luiz Nunes, Neele Engelmann, Alex Wiegmann, Marcelo de Araújo

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref Exploring the psychology of LLMs' moral and legal reasoning. Artificial Intelligence, Volume 224, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11273 2024-05-21 cs.AI cs.CL cs.CV cs.MM 62%

Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

Yunxin Li, Shenyuan Jiang, Baotian Hu, Longyue Wang, Wanqi Zhong, Wenhan Luo, Lin Ma, Min Zhang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 22 pages, 13 figures. Project Website: https://uni-moe.github.io/. Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏