arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1836 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1836 篇

2305.04547 2023-05-09 cs.CL 57%

Diffusion Theory as a Scalpel: Detecting and Purifying Poisonous Dimensions in Pre-trained Language Models Caused by Backdoor or Bias

Zhiyuan Zhang, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

Comments Accepted by Findings of ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08001 2023-02-17 cs.AI cs.MA 57%

Learning Density-Based Correlated Equilibria for Markov Games

Libo Zhang, Yang Chen, Toru Takisaka, Bakh Khoussainov, Michael Witbrock, Jiamou Liu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12566 2023-01-31 cs.CL cs.IR 57%

Improving Cross-lingual Information Retrieval on Low-Resource Languages via Optimal Transport Distillation

Zhiqi Huang, Puxuan Yu, James Allan

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11737 2023-01-30 cs.LG 57%

Modeling human road crossing decisions as reward maximization with visual perception limitations

Yueyang Wang, Aravinda Ramakrishnan Srinivasan, Jussi P. P. Jokinen, Antti Oulasvirta, Gustav Markkula

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 6 pages, 5 figures,1 table, manuscript created for consideration at IEEE IV 2023 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08094 2023-01-05 cs.CV cs.LG 57%

LIMEcraft: Handcrafted superpixel selection and inspection for Visual eXplanations

Weronika Hryniewska, Adrianna Grudzień, Przemysław Biecek

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Journal ref Machine Learning (2022) 1-18

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06833 2022-11-30 cs.CY 57%

Beyond Ads: Sequential Decision-Making Algorithms in Law and Public Policy

Peter Henderson, Ben Chugg, Brandon Anderson, Daniel E. Ho

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Version 1 presented at Causal Inference Challenges in Sequential Decision Making: Bridging Theory and Practice (2021), a NeurIPS 2021 Workshop; Version 2 presented at the 2nd ACM Symposium on Computer Science and Law (2022) (DOI: https://dl.acm.org/doi/10.1145/3511265.3550439)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04491 2022-10-11 cs.LG stat.ML 57%

A survey of Identification and mitigation of Machine Learning algorithmic biases in Image Analysis

Laurent Risser, Agustin Picard, Lucas Hervier, Jean-Michel Loubes

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10087 2022-08-23 cs.CY 57%

A Trust Framework for Government Use of Artificial Intelligence and Automated Decision Making

Pia Andrews, Tim de Sousa, Bruce Haefele, Matt Beard, Marcus Wigan, Abhinav Palia, Kathy Reid, Saket Narayan, Morgan Dumitru, Alex Morrison, Geoff Mason, Aurelie Jacquet

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Comments Comments were integrated into the paper from all peer reviewers. Am happy to provide a copied history of comments if useful

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02725 2022-08-23 cs.IR cs.CL 57%

Match-Prompt: Improving Multi-task Generalization Ability for Neural Text Matching via Prompt Learning

Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted by CIKM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08790 2022-08-19 cs.AI 57%

Explainable Reinforcement Learning on Financial Stock Trading using SHAP

Satyam Kumar, Mendhikar Vishal, Vadlamani Ravi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 28 pages; 3 Tables; 21 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.07120 2022-07-27 cs.LG cs.CV 57%

Making Corgis Important for Honeycomb Classification: Adversarial Attacks on Concept-based Explainability Tools

Davis Brown, Henry Kvinge

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments AdvML Frontiers 2022 @ ICML 2022 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.00474 2022-06-02 cs.AI cs.HC 57%

Towards Responsible AI: A Design Space Exploration of Human-Centered Artificial Intelligence User Interfaces to Investigate Fairness

Yuri Nakao, Lorenzo Strappelli, Simone Stumpf, Aisha Naseer, Daniele Regoli, Giulia Del Gamba

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 44 pages, 17 figures, the draft of a paper on International Journal of Human-Computer Interaction

Journal ref International Journal of Human-Computer Interaction, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03824 2022-05-10 cs.AI 57%

A Survey on AI Sustainability: Emerging Trends on Learning Algorithms and Research Challenges

Zhenghua Chen, Min Wu, Alvin Chan, Xiaoli Li, Yew-Soon Ong

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00656 2022-05-03 cs.CL 57%

Debiased Contrastive Learning of Unsupervised Sentence Representations

Kun Zhou, Beichen Zhang, Wayne Xin Zhao, Ji-Rong Wen

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments 11 pages, accepted by ACL 2022 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09622 2022-04-21 cs.HC cs.GL cs.LG 57%

A Brief Guide to Designing and Evaluating Human-Centered Interactive Machine Learning

Kory W. Mathewson, Patrick M. Pilarski

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 7 pages, 1 figure, Published at ML Evaluation Standards Workshop at ICLR 2022. arXiv admin note: substantial text overlap with arXiv:1905.06289

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.06933 2022-04-05 cs.LG stat.ML 57%

The reinforcement learning-based multi-agent cooperative approach for the adaptive speed regulation on a metallurgical pickling line

Anna Bogomolova, Kseniia Kingsep, Boris Voskresenskii

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.02038 2022-03-23 cs.SE cs.LG 57%

Fair-SSL: Building fair ML Software with less data

Joymallya Chakraborty, Suvodeep Majumder, Huy Tu

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

Journal ref International Workshop on Equitable Data and Technology (FairWare 2022 ), May 9, 2022, Pittsburgh, PA, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09616 2022-03-21 astro-ph.CO cs.LG 57%

DeepLSS: breaking parameter degeneracies in large scale structure with deep learning analysis of combined probes

Tomasz Kacprzak, Janis Fluri

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments 18 pages, 10 figures, 2 tables, submitted to Physical Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01252 2022-03-11 cs.CY 57%

Australia's Approach to AI Governance in Security and Defence

Susannah Kate Devitt, Damian Copeland

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Comments 60 pages, 7 boxes, 2 figures, 2 annexes, submitted for Eds M. Raska, Z. Stanley-Lockman, & R. Bitzinger. AI Governance for National Security and Defence: Assessing Military AI Strategic Perspectives. Routledge

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.11629 2022-02-24 cs.AI stat.ML 57%

A Complete Criterion for Value of Information in Soluble Influence Diagrams

Chris van Merwijk, Ryan Carey, Tom Everitt

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments In Proceedings of the AAAI 2022 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.15208 2022-02-04 cs.CV cs.AI 57%

HRNET: AI on Edge for mask detection and social distancing

Kinshuk Sengupta, Praveen Ranjan Srivastava

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Journal ref SN Computer Science, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.01659 2022-01-06 cs.CY 57%

From the Ground Truth Up: Doing AI Ethics from Practice to Principles

James Brusseau

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

Journal ref AI & Soc (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.15234 2022-01-05 cs.AI cs.GT 57%

Artificial Intelligence Development Races in Heterogeneous Settings

Theodor Cimpeanu, Francisco C. Santos, Luis Moniz Pereira, Tom Lenaerts, The Anh Han

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 42 pages, 13 figures, accepted for publication in Nature Scientific Reports

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06883 2021-12-14 cs.SE cs.AI cs.DC 57%

A Methodology for a Scalable, Collaborative, and Resource-Efficient Platform to Facilitate Healthcare AI Research

Raphael Y. Cohen, Vesela P. Kovacheva

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.16122 2021-11-03 cs.HC cs.CY 57%

Zombies in the Loop? Humans Trust Untrustworthy AI-Advisors for Ethical Decisions

Sebastian Krügel, Andreas Ostermaier, Matthias Uhl

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.05164 2021-10-12 cs.CY 57%

Ethical Assurance: A practical approach to the responsible design, development, and deployment of data-driven technologies

Christopher Burr, David Leslie

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.00421 2021-09-22 cs.CL 57%

The Highs and Lows of Simple Lexical Domain Adaptation Approaches for Neural Machine Translation

Nikolay Bogoychev, Pinzhen Chen

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted at Workshop on Insights from Negative Results in NLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.05388 2021-09-14 cs.CL 57%

The Impact of Positional Encodings on Multilingual Compression

Vinit Ravishankar, Anders Søgaard

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.14099 2021-07-30 cs.CY 57%

The ghost of AI governance past, present and future: AI governance in the European Union

Charlotte Stix

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.13076 2021-07-29 cs.HC cs.LG 57%

Interactive Storytelling for Children: A Case-study of Design and Development Considerations for Ethical Conversational AI

ennifer Chubba, Sondess Missaouib, Shauna Concannonc, Liam Maloneyb, James Alfred Walker

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏