arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1836 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1836 篇

2402.08797 2024-02-15 cs.CY 57%

Computing Power and the Governance of Artificial Intelligence

Girish Sastry, Lennart Heim, Haydn Belfield, Markus Anderljung, Miles Brundage, Julian Hazell, Cullen O'Keefe, Gillian K. Hadfield, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Janet Egan, Robert F. Trager, Shahar Avin, Adrian Weller, Yoshua Bengio, Diane Coyle

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Figures can be accessed at: https://github.com/lheim/CPGAI-Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04017 2024-02-13 cs.LG q-bio.QM 57%

PGraphDTA: Improving Drug Target Interaction Prediction using Protein Language Models and Contact Maps

Rakesh Bal, Yijia Xiao, Wei Wang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments AI for Science Workshop, NeurIPS 2023. 11 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05786 2024-02-12 cs.AI cs.GT 57%

Prompting Fairness: Artificial Intelligence as Game Players

Jazmia Henry

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00745 2024-02-02 cs.CL 57%

Enhancing Ethical Explanations of Large Language Models through Iterative Symbolic Refinement

Xin Quan, Marco Valentino, Louise A. Dennis, André Freitas

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Camera-ready for EACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15798 2024-01-30 cs.CL 57%

UnMASKed: Quantifying Gender Biases in Masked Language Models through Linguistically Informed Job Market Prompts

Iñigo Parra

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments EACL 2024 SRW

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09852 2024-01-19 cs.CV cs.AI 57%

Enhancing the Fairness and Performance of Edge Cameras with Explainable AI

Truong Thanh Hung Nguyen, Vo Thanh Khang Nguyen, Quoc Hung Cao, Van Binh Truong, Quoc Khanh Nguyen, Hung Cao

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments IEEE ICCE 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18007 2023-12-01 astro-ph.IM astro-ph.GA cs.LG 57%

Towards out-of-distribution generalization in large-scale astronomical surveys: robust networks learn similar representations

Yash Gondhalekar, Sultan Hassan, Naomi Saphra, Sambatra Andrianomena

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments Accepted to Machine Learning and the Physical Sciences Workshop, NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04955 2023-11-17 cs.LG 57%

Information-Theoretic Bounds on The Removal of Attribute-Specific Bias From Neural Networks

Jiazhi Li, Mahyar Khayatkhoei, Jiageng Zhu, Hanchen Xie, Mohamed E. Hussein, Wael AbdAlmageed

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

Comments 15 pages, 4 figures, 3 tables. To appear in Algorithmic Fairness through the Lens of Time Workshop at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09687 2023-11-17 cs.CL 57%

Inducing Political Bias Allows Language Models Anticipate Partisan Reactions to Controversies

Zihao He, Siyi Guo, Ashwin Rao, Kristina Lerman

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16607 2023-11-14 cs.CL 57%

On the Interplay between Fairness and Explainability

Stephanie Brandl, Emanuele Bugliarello, Ilias Chalkidis

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL

Comments 15 pages (incl Appendix), 4 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18130 2023-11-09 cs.CL cs.HC 57%

DELPHI: Data for Evaluating LLMs' Performance in Handling Controversial Issues

David Q. Sun, Artem Abzaliev, Hadas Kotek, Zidi Xiu, Christopher Klein, Jason D. Williams

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

Comments Accepted to EMNLP Industry Track 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04411 2023-11-08 cs.LG 57%

Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL

Yang Yue, Rui Lu, Bingyi Kang, Shiji Song, Gao Huang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments 31 pages, 20 figures

Journal ref NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01550 2023-11-06 cs.AI econ.GN q-fin.EC 57%

Market Concentration Implications of Foundation Models

Jai Vipra, Anton Korinek

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Working Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11090 2023-11-01 cs.CV cs.LG stat.AP 57%

Fairness Explainability using Optimal Transport with Applications in Image Classification

Philipp Ratz, François Hu, Arthur Charpentier

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18832 2023-10-31 cs.AI 57%

Responsible AI (RAI) Games and Ensembles

Yash Gupta, Runtian Zhai, Arun Suggala, Pradeep Ravikumar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16360 2023-10-26 cs.AI cs.RO 57%

A Comprehensive Review of AI-enabled Unmanned Aerial Vehicle: Trends, Vision , and Challenges

Osim Kumar Pal, Md Sakib Hossain Shovon, M. F. Mridha, Jungpil Shin

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09217 2023-10-16 cs.AI 57%

Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI

Jason Hausenloy, Andrea Miotti, Claire Dennis

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05876 2023-10-10 cs.AI 57%

AI Systems of Concern

Kayla Matteucci, Shahar Avin, Fazl Barez, Seán Ó hÉigeartaigh

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 9 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11691 2023-09-22 cs.AI cs.CR 57%

RAI4IoE: Responsible AI for Enabling the Internet of Energy

Minhui Xue, Surya Nepal, Ling Liu, Subbu Sethuvenkatraman, Xingliang Yuan, Carsten Rudolph, Ruoxi Sun, Greg Eisenhauer

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Accepted to IEEE International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15514 2023-09-12 cs.AI 57%

International Governance of Civilian AI: A Jurisdictional Certification Approach

Robert Trager, Ben Harack, Anka Reuel, Allison Carnegie, Lennart Heim, Lewis Ho, Sarah Kreps, Ranjit Lall, Owen Larter, Seán Ó hÉigeartaigh, Simon Staffell, José Jaime Villalobos

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06326 2023-09-07 cs.CY 57%

Bad, mad, and cooked: Moral responsibility for civilian harms in human-AI military teams

Susannah Kate Devitt

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 30 pages, accepted for publication in Jan Maarten Schraagen (Ed.) 'Responsible Use of AI in Military Systems', CRC Press [Forthcoming]

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02996 2023-08-08 cs.NI cs.CY cs.DC 57%

A Review of Gaps between Web 4.0 and Web 3.0 Intelligent Network Infrastructure

Zihan Zhou, Zihao Li, Xiaoshuai Zhang, Yunqing Sun, Hao Xu

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Comments 6 pages 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02025 2023-08-07 cs.CY cs.IR 57%

Applications and Societal Implications of Artificial Intelligence in Manufacturing: A Systematic Review

John P. Nelson, Justin B. Biddle, Philip Shapira

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14672 2023-06-27 stat.ML cs.LG 57%

PWSHAP: A Path-Wise Explanation Model for Targeted Variables

Lucile Ter-Minassian, Oscar Clivio, Karla Diaz-Ordaz, Robin J. Evans, Chris Holmes

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Journal ref International Conference on Machine Learning 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08959 2023-06-16 cs.CY 57%

Statutory Professions in AI governance and their consequences for explainable AI

Labhaoise NiFhaolain, Andrew Hines, Vivek Nallur

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Accepted for publication at xAI-2023 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10743 2023-06-13 cs.CY cs.HC 57%

The Ethics of AI-Generated Maps: A Study of DALLE 2 and Implications for Cartography

Yuhao Kang, Qianheng Zhang, Robert Roth

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Comments 9 pages, 3 figures, GIScience 2023 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.00796 2023-06-01 cs.LG cs.IT math.IT 57%

On Balancing Bias and Variance in Unsupervised Multi-Source-Free Domain Adaptation

Maohao Shen, Yuheng Bu, Gregory Wornell

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments ICML 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11837 2023-05-26 cs.SE cs.AI 57%

Comparing Software Developers with ChatGPT: An Empirical Investigation

Nathalia Nascimento, Paulo Alencar, Donald Cowan

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14865 2023-05-25 cs.GT cs.CY 57%

A Game-Theoretic Framework for AI Governance

Na Zhang, Kun Yue, Chao Fang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03870 2023-05-11 cs.LG 57%

Knowledge Transfer from Teachers to Learners in Growing-Batch Reinforcement Learning

Patrick Emedom-Nnamdi, Abram L. Friesen, Bobak Shahriari, Nando de Freitas, Matt W. Hoffman

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments Reincarnating Reinforcement Learning Workshop at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏