arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9311 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9311 篇

2209.06529 2022-09-15 cs.LG cs.CR 79%

Data Privacy and Trustworthy Machine Learning

Martin Strobel, Reza Shokri

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments Copyright ©2022, IEEE

Journal ref Published in: IEEE Security & Privacy ( Volume: 20, Issue: 5, Sept.-Oct. 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12615 2022-07-27 cs.LG 79%

Exploring the Design of Adaptation Protocols for Improved Generalization and Machine Learning Safety

Puja Trivedi, Danai Koutra, Jayaraman J. Thiagarajan

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments Principles of Distribution Shift (PODS) Workshop at ICML 2022, 4 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10809 2022-07-25 cs.CR cs.AI 79%

Security and Safety Aspects of AI in Industry Applications

Hans Dermot Doran

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments As presented at the Embedded World Conference, Nuremberg, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.09312 2022-07-20 eess.IV cs.CV cs.LG 79%

Towards Trustworthy Healthcare AI: Attention-Based Feature Learning for COVID-19 Screening With Chest Radiography

Kai Ma, Pengcheng Xi, Karim Habashy, Ashkan Ebadi, Stéphane Tremblay, Alexander Wong

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments Accepted to 39th International Conference on Machine Learning, Workshop on Healthcare AI and COVID-19

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.07505 2022-07-15 eess.SY cs.LG cs.SY 79%

Closing the Loop: A Framework for Trustworthy Machine Learning in Power Systems

Jochen Stiasny, Samuel Chevalier, Rahul Nellikkath, Brynjar Sævarsson, Spyros Chatzivasileiadis

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments In proceedings of the 11th Bulk Power Systems Dynamics and Control Symposium (IREP 2022), July 25-30, 2022, Banff, Canada. 21 pages, 12 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09523 2022-06-22 eess.AS cs.CY cs.SD 79%

Towards Trustworthy Edge Intelligence: Insights from Voice-Activated Services

W. T. Hutiri, A. Y. Ding

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.03044 2022-06-22 cs.AI cs.LO cs.NE cs.SE 79%

CAISAR: A platform for Characterizing Artificial Intelligence Safety and Robustness

Julien Girard-Satabin, Michele Alberti, François Bobot, Zakaria Chihani, Augustin Lemesle

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Journal ref AISafety, Jul 2022, Vienne, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07506 2022-06-16 cs.HC cs.CY 79%

Legal Provocations for HCI in the Design and Development of Trustworthy Autonomous Systems

Lachlan D. Urquhart, Glenn McGarry, Andy Crabtree

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.01909 2022-06-07 cs.LG 79%

Toward Learning Robust and Invariant Representations with Alignment Regularization and Data Augmentation

Haohan Wang, Zeyi Huang, Xindi Wu, Eric P. Xing

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

Comments to appear at KDD 2022, the software package is at https://github.com/jyanln/AlignReg. arXiv admin note: text overlap with arXiv:2011.13052

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.14095 2022-05-31 cs.CV cs.AI 79%

PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining

Yuting Gao, Jinfeng Liu, Zihan Xu, Jun Zhang, Ke Li, Rongrong Ji, Chunhua Shen

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.15059 2022-05-09 cs.AI 79%

A Benchmark and Comprehensive Survey on Knowledge Graph Entity Alignment via Representation Learning

Rui Zhang, Bayu Distiawan Trisedy, Miao Li, Yong Jiang, Jianzhong Qi

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments to appear in VLDB Journal, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.09234 2022-03-28 cs.LG 79%

Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model Behavior

Angie Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik Strobelt

专题命中 安全评测 :alignment(title);trustworthy(abstract);分类 cs.LG

Comments 17 pages, 10 figures. Published in CHI 2022. For more details, see http://shared-interest.csail.mit.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12653 2022-02-28 cs.LG stat.ML 79%

Bayesian autoencoders with uncertainty quantification: Towards trustworthy anomaly detection

Bang Xiang Yong, Alexandra Brintrup

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.03042 2022-02-15 cs.LG cs.CR 79%

Towards a Robust and Trustworthy Machine Learning System Development: An Engineering Perspective

Pulei Xiong, Scott Buffett, Shahrear Iqbal, Philippe Lamontagne, Mohammad Mamun, Heather Molyneaux

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments 20 pages (58 pages pre-print), 6 figures

Journal ref Journal of Information Security and Applications 65 (2022) 103121

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.05313 2022-02-14 cs.AI cs.SE 79%

Integrating Testing and Operation-related Quantitative Evidences in Assurance Cases to Argue Safety of Data-Driven AI/ML Components

Michael Kläs, Lisa Jöckel, Rasmus Adler, Jan Reich

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.01934 2022-02-07 cs.LG 79%

Smartphone-based Hard-braking Event Detection at Scale for Road Safety Services

Luyang Liu, David Racz, Kara Vaillancourt, Julie Michelman, Matt Barnes, Stefan Mellem, Paul Eastham, Bradley Green, Charles Armstrong, Rishi Bal, Shawn O'Banion, Feng Guo

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.09998 2021-12-21 cs.RO cs.LG cs.SY eess.SY 79%

Learning-based methods to model small body gravity fields for proximity operations: Safety and Robustness

Daniel Neamati, Yashwanth Kumar Nakka, Soon-Jo Chung

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments Accepted Scitech, AI for Space

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03024 2021-12-07 cs.CL 79%

Domain-oriented Language Pre-training with Adaptive Hybrid Masking and Optimal Transport Alignment

Denghui Zhang, Zixuan Yuan, Yanchi Liu, Hao Liu, Fuzhen Zhuang, Hui Xiong, Haifeng Chen

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02707 2021-10-07 cs.SE cs.AI 79%

Trustworthy Artificial Intelligence and Process Mining: Challenges and Opportunities

Andrew Pery, Majid Rafiei, Michael Simon, Wil M. P. van der Aalst

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01232 2021-10-05 cs.AI 79%

Benchmarking Safety Monitors for Image Classifiers with Machine Learning

Raul Sena Ferreira, Jean Arlat, Jeremie Guiochet, Hélène Waeselynck

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Journal ref 26th IEEE Pacific Rim International Symposium on Dependable Computing (PRDC 2021), IEEE, Dec 2021, Perth, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.10142 2021-09-28 cs.CV cs.LG 79%

Safety Metrics for Semantic Segmentation in Autonomous Driving

Chih-Hong Cheng, Alois Knoll, Hsuan-Cheng Liao

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments Paper accepted at IEEE AI Test'21

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.01287 2021-09-16 math.OC cs.LG 79%

Safety Verification and Robustness Analysis of Neural Networks via Quadratic Constraints and Semidefinite Programming

Mahyar Fazlyab, Manfred Morari, George J. Pappas

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01531 2021-09-06 cs.LG 79%

MACEst: The reliable and trustworthy Model Agnostic Confidence Estimator

Rhys Green, Matthew Rowe, Alberto Polleri

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06080 2021-08-16 cs.LG 79%

TDM: Trustworthy Decision-Making via Interpretability Enhancement

Daoming Lyu, Fangkai Yang, Hugh Kwon, Wen Dong, Levent Yilmaz, Bo Liu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Journal ref IEEE Transactions on Emerging Topics in Computational Intelligence 0 (2021) 1-12

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.08231 2021-08-13 cs.CL 79%

Word Alignment by Fine-tuning Embeddings on Parallel Corpora

Zi-Yi Dou, Graham Neubig

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments EACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.06466 2021-06-08 cs.LG stat.ML 79%

How Interpretable and Trustworthy are GAMs?

Chun-Hao Chang, Sarah Tan, Ben Lengerich, Anna Goldenberg, Rich Caruana

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments Accepted in 2021 KDD

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.00512 2021-06-02 cs.LG 79%

The Care Label Concept: A Certification Suite for Trustworthy and Resource-Aware Machine Learning

Katharina Morik, Helena Kotthaus, Lukas Heppe, Danny Heinrich, Raphael Fischer, Andreas Pauly, Nico Piatkowski

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.05434 2021-04-26 cs.SE cs.AI 79%

Developing and Operating Artificial Intelligence Models in Trustworthy Autonomous Systems

Silverio Martínez-Fernández, Xavier Franch, Andreas Jedlitschka, Marc Oriol, Adam Trendowicz

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 9 pages, 1 figure, preprint. Accepted in RCIS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.13073 2020-07-21 cs.CV cs.LG eess.IV 79%

Attributional Robustness Training using Input-Gradient Spatial Alignment

Mayank Singh, Nupur Kumari, Puneet Mangla, Abhishek Sinha, Vineeth N Balasubramanian, Balaji Krishnamurthy

专题命中 安全评测 :alignment(title);trustworthy(abstract);分类 cs.LG

Comments ECCV 2020, Code at https://github.com/nupurkmr9/Attributional-Robustness

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.01155 2020-07-01 cs.SE cs.LG 79%

Towards Probability-based Safety Verification of Systems with Components from Machine Learning

Hermann Kaindl, Stefan Kramer

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments Second (revised) version for public access

详情

展开后加载摘要…

URL PDF HTML 收藏