Mediating Artificial Intelligence Developments through Negative and Positive Incentives
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY
Comments To appear in the 2022 Springer book "The Future Circle of Healthcare: AI, 3D Printing, Longevity, Ethics, and Uncertainty Mitigation" edited by Sepehr Ehsani, Patrick Glauner, Philipp Plugmann and Florian M. Thieringer
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
Comments 16 pages, 3 figures, pre-print for the International Workshop on Coordination, Organizations, Institutions, Norms and Ethics for Governance of Multi-Agent Systems (COINE), co-located with AAMAS 2020
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
Comments To be published in ACT Transactions on Cyber-Physical Systems Special Issue on Artificial Intelligence and Cyber-Physical Systems. arXiv admin note: text overlap with arXiv:2009.00738
专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG
专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG
Journal ref Frontiers in Robotics and AI, 2021
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY
Journal ref Sci Eng Ethics 27, 15 (2021)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY
Comments 4 pages, 3 Figures, Policy Briefing, Based on research funded by EPSR and carried out by UCL STEaPP NIPC-ALIoTT collaboration project under the PETRAS Cybersecurity Hub
Journal ref The Internet of Things in Ports: Six Key Security and Governance Challenges for the UK (A Policy Brief). London: PETRAS National Centre of Excellence for IoT System Cybersecurity (2019)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY
Comments Accepted to NeurIPS 2020 Workshop on Dataset Curation and Security; Also accepted at Navigating the Broader Impacts of AI Research Workshop. All authors contributed equally. The list of authors is arranged alphabetically
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY
Comments 12 pages, 10 figure
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG
Comments Accepted to ICML2020. Code is available at https://github.com/bfshi/InfoDrop
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG
专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG
Comments To appear at the 23rd Design, Automation and Test in Europe (DATE 2020). Grenoble, France
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG
Comments IEEE Transactions on Signal Processing, Early Access, Mar. 2020
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG
Comments 6 pages
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
Comments 10 pages
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI
Comments 28 pages
专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG
Comments 18 pages, 9 figures
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
Comments 7 pages, 1 figure, submitted
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
Comments Presented at the 1st International Workshop on AI and Ethics, Sunday 25th January 2015, Hill Country A, Hyatt Regency Austin. Will appear in the workshop proceedings published by AAAI
机构 * Intelligent Systems Program University of Pittsburgh Pittsburgh Pennsylvania USA ; Intelligent Systems Program University of Pittsburgh
专题命中 AI治理与伦理 :分类 cs.CL、cs.AI、cs.LG;safety(comments)
Comments 13 pages, 2 figures, 2nd ConventicLe on Artificial Intelligence Regulation and Safety Workshop at ICAIL 2025
专题命中 AI治理与伦理 :trustworthy(abstract,comments)
Comments To be published in Second International Symposium on Trustworthy Autonomous Systems, September 15 18, 2024, Austin, Texas
专题命中 AI治理与伦理 :分类 cs.CL、cs.CY、cs.LG;trustworthy(comments)
Comments Accepted for publication at Third Workshop on Trustworthy Natural Language Processing, ACL 2023
专题命中 AI治理与伦理 :分类 cs.CL、cs.CY、cs.LG;trustworthy(comments)
Comments Accepted to 2022 Trustworthy and Socially Responsible Machine Learning (TSRML 2022) Workshop at NeurIPS 2022