arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1836 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1836 篇

2010.00403 2021-06-09 cs.AI cs.MA math.DS nlin.AO q-bio.PE 57%

Mediating Artificial Intelligence Developments through Negative and Positive Incentives

The Anh Han, Luis Moniz Pereira, Tom Lenaerts, Francisco C. Santos

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.15133 2021-06-01 cs.CY 57%

An Assessment of the AI Regulation Proposed by the European Commission

Patrick Glauner

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments To appear in the 2022 Springer book "The Future Circle of Healthcare: AI, 3D Printing, Longevity, Ethics, and Uncertainty Mitigation" edited by Sepehr Ehsani, Patrick Glauner, Philipp Plugmann and Florian M. Thieringer

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.01056 2021-05-17 cs.AI cs.MA 57%

Improving Confidence in the Estimation of Values and Norms

Luciano Cavalcante Siebert, Rijk Mercuur, Virginia Dignum, Jeroen van den Hoven, Catholijn Jonker

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 16 pages, 3 figures, pre-print for the International Workshop on Coordination, Organizations, Institutions, Norms and Ethics for Governance of Multi-Agent Systems (COINE), co-located with AAMAS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.02851 2021-05-07 cs.AI 57%

Algorithmic Ethics: Formalization and Verification of Autonomous Vehicle Obligations

Colin Shea-Blymyer, Houssam Abbas

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments To be published in ACT Transactions on Cyber-Physical Systems Special Issue on Artificial Intelligence and Cyber-Physical Systems. arXiv admin note: text overlap with arXiv:2009.00738

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.03103 2021-03-05 cs.LG 57%

Interpretable Artificial Intelligence through the Lens of Feature Interaction

Michael Tsang, James Enouen, Yan Liu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.02647 2021-03-04 cs.RO cs.CV cs.LG 57%

From Learning to Relearning: A Framework for Diminishing Bias in Social Robot Navigation

Juana Valeria Hurtado, Laura Londoño, Abhinav Valada

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Journal ref Frontiers in Robotics and AI, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.12406 2021-02-25 cs.CY 57%

Actionable Principles for Artificial Intelligence Policy: Three Pathways

Charlotte Stix

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Journal ref Sci Eng Ethics 27, 15 (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.08812 2021-01-25 cs.CY 57%

The Internet of Things in Ports: Six Key Security and Governance Challenges for the UK (Policy Brief)

Feja Lesniewska, Uchenna D Ani, Jeremy M Watson, Madeline Carr

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 4 pages, 3 Figures, Policy Briefing, Based on research funded by EPSR and carried out by UCL STEaPP NIPC-ALIoTT collaboration project under the PETRAS Cybersecurity Hub

Journal ref The Internet of Things in Ports: Six Key Security and Governance Challenges for the UK (A Policy Brief). London: PETRAS National Centre of Excellence for IoT System Cybersecurity (2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.02048 2020-12-04 cs.CY 57%

Ethical Testing in the Real World: Evaluating Physical Testing of Adversarial Machine Learning

Kendra Albert, Maggie Delano, Jonathon Penney, Afsaneh Rigot, Ram Shankar Siva Kumar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Accepted to NeurIPS 2020 Workshop on Dataset Curation and Security; Also accepted at Navigating the Broader Impacts of AI Research Workshop. All authors contributed equally. The list of authors is arranged alphabetically

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.12530 2020-08-31 cs.CY cs.SY eess.SY physics.soc-ph 57%

Investigating Taxi and Uber competition in New York City: Multi-agent modeling by reinforcement-learning

Saeed Vasebi, Yeganeh M. Hayeri

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 12 pages, 10 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.04254 2020-08-11 cs.LG cs.CV stat.ML 57%

Informative Dropout for Robust Representation Learning: A Shape-bias Perspective

Baifeng Shi, Dinghuai Zhang, Qi Dai, Zhanxing Zhu, Yadong Mu, Jingdong Wang

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

Comments Accepted to ICML2020. Code is available at https://github.com/bfshi/InfoDrop

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.01522 2020-07-06 cs.LG stat.ML 57%

Dueling Deep Q-Network for Unsupervised Inter-frame Eye Movement Correction in Optical Coherence Tomography Volumes

Yasmeen M. George, Suman Sedai, Bhavna J. Antony, Hiroshi Ishikawa, Gadi Wollstein, Joel S. Schuman, Rahil Garnavi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.01978 2020-05-18 cs.LG stat.ML 57%

FANNet: Formal Analysis of Noise Tolerance, Training Bias and Input Sensitivity in Neural Networks

Mahum Naseer, Mishal Fatima Minhas, Faiq Khalid, Muhammad Abdullah Hanif, Osman Hasan, Muhammad Shafique

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments To appear at the 23rd Design, Automation and Test in Europe (DATE 2020). Grenoble, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.04644 2020-04-10 cs.LG 57%

On the Ethics of Building AI in a Responsible Manner

Shai Shalev-Shwartz, Shaked Shammah, Amnon Shashua

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.00844 2020-04-08 cs.DC cs.IT cs.LG math.IT 57%

Machine Learning at the Wireless Edge: Distributed Stochastic Gradient Descent Over-the-Air

Mohammad Mohammadi Amiri, Deniz Gunduz

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments IEEE Transactions on Signal Processing, Early Access, Mar. 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.11157 2020-03-26 cs.CY 57%

AI loyalty: A New Paradigm for Aligning Stakeholder Interests

Anthony Aguirre, Gaia Dempsey, Harry Surden, Peter B. Reiner

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.10985 2020-02-04 cs.AI 57%

AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence

Jeff Clune

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.01723 2019-07-04 stat.ML cs.LG stat.AP 57%

Towards Interpretable Deep Extreme Multi-label Learning

Yihuang Kang, I-Ling Cheng, Wenjui Mao, Bowen Kuo, Pei-Ju Lee

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.04585 2019-02-26 cs.AI q-fin.GN stat.ML 57%

Categorizing Variants of Goodhart's Law

David Manheim, Scott Garrabrant

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.08221 2019-01-25 cs.AI 57%

When is it right and good for an intelligent autonomous vehicle to take over control (and hand it back)?

Ajit Narayanan

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.06064 2018-07-18 cs.LG stat.ML 57%

Online Robust Policy Learning in the Presence of Unknown Adversaries

Aaron J. Havens, Zhanhong Jiang, Soumik Sarkar

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 18 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.10322 2018-06-28 cs.AI 57%

The Virtuous Machine - Old Ethics for New Technology?

Nicolas Berberich, Klaus Diepold

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.10820 2018-05-29 cs.AI 57%

Local Rule-Based Explanations of Black Box Decision Systems

Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Dino Pedreschi, Franco Turini, Fosca Giannotti

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.09854 2018-01-31 cs.AI 57%

Algorithms for the Greater Good! On Mental Modeling and Acceptable Symbiosis in Human-AI Collaboration

Tathagata Chakraborti, Subbarao Kambhampati

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.06035 2017-11-17 cs.AI 57%

From Algorithmic Black Boxes to Adaptive White Boxes: Declarative Decision-Theoretic Ethical Programs as Codes of Ethics

Martijn van Otterlo

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 7 pages, 1 figure, submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
1504.03592 2015-04-15 cs.AI 57%

Towards Verifiably Ethical Robot Behaviour

Louise A. Dennis, Michael Fisher, Alan F. T. Winfield

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Presented at the 1st International Workshop on AI and Ethics, Sunday 25th January 2015, Hill Country A, Hyatt Regency Austin. Will appear in the workshop proceedings published by AAAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02992 2025-10-27 cs.AI cs.CL cs.LG 56%

Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation

Li Zhang, Kevin D. Ashley

机构 * Intelligent Systems Program University of Pittsburgh Pittsburgh Pennsylvania USA Intelligent Systems Program University of Pittsburgh

专题命中 AI治理与伦理 :分类 cs.CL、cs.AI、cs.LG;safety(comments)

Comments 13 pages, 2 figures, 2nd ConventicLe on Artificial Intelligence Regulation and Safety Workshop at ICAIL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02386 2024-08-06 cs.HC 56%

Responsibility and Regulation: Exploring Social Measures of Trust in Medical AI

Glenn McGarry, Andy Crabtree, Lachlan Urquhart, Alan Chamberlain

专题命中 AI治理与伦理 :trustworthy(abstract,comments)

Comments To be published in Second International Symposium on Trustworthy Autonomous Systems, September 15 18, 2024, Austin, Texas

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05307 2023-06-09 cs.CL cs.CY cs.LG stat.ML 56%

Are fairness metric scores enough to assess discrimination biases in machine learning?

Fanny Jourdan, Laurent Risser, Jean-Michel Loubes, Nicholas Asher

专题命中 AI治理与伦理 :分类 cs.CL、cs.CY、cs.LG;trustworthy(comments)

Comments Accepted for publication at Third Workshop on Trustworthy Natural Language Processing, ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14402 2022-11-29 cs.CL cs.CY cs.LG 56%

An Analysis of Social Biases Present in BERT Variants Across Multiple Languages

Aristides Milios, Parishad BehnamGhader

专题命中 AI治理与伦理 :分类 cs.CL、cs.CY、cs.LG;trustworthy(comments)

Comments Accepted to 2022 Trustworthy and Socially Responsible Machine Learning (TSRML 2022) Workshop at NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏