arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9248 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9248 篇

2310.12443 2023-10-20 cs.IR cs.AI cs.CL 81%

Know Where to Go: Make LLM a Relevant, Responsible, and Trustworthy Searcher

Xiang Shi, Jiawei Liu, Yinpeng Liu, Qikai Cheng, Wei Lu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments 14 pages, 4 figures, under peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08215 2023-10-13 cs.LG cs.AI 81%

Trustworthy Machine Learning

Bálint Mucsányi, Michael Kirchhof, Elisa Nguyen, Alexander Rubinstein, Seong Joon Oh

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 373 pages, textbook at the University of Tübingen

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12324 2023-09-25 cs.CY cs.LG cs.SY eess.SY 81%

Aviation Safety Risk Analysis and Flight Technology Assessment Issues

Shuanghe Liu

专题命中 安全评测 :safety(title,abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09450 2023-09-19 cs.CY cs.AI cs.HC 81%

Are You Worthy of My Trust?: A Socioethical Perspective on the Impacts of Trustworthy AI Systems on the Environment and Human Society

Jamell Dacon

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12315 2023-08-31 cs.LG cs.AI 81%

Trustworthy Representation Learning Across Domains

Ronghang Zhu, Dongliang Guo, Daiqing Qi, Zhixuan Chu, Xiang Yu, Sheng Li

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 38 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04445 2023-08-10 cs.LG cs.AI 81%

Getting from Generative AI to Trustworthy AI: What LLMs might learn from Cyc

Doug Lenat, Gary Marcus

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 21 pages, 1 Figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00076 2023-08-02 cs.AI cs.LG stat.ML 81%

Crowd Safety Manager: Towards Data-Driven Active Decision Support for Planning and Control of Crowd Events

Panchamy Krishnakumari, Sascha Hoogendoorn-Lanser, Jeroen Steenbakkers, Serge Hoogendoorn

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments Submitted to TRB Annual Meeting 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16851 2023-08-01 cs.LG cs.AI 81%

Towards Trustworthy and Aligned Machine Learning: A Data-centric Survey with Causality Perspectives

Haoyang Liu, Maheep Chaudhary, Haohan Wang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 47 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04683 2023-07-11 cs.CL cs.AI 81%

CORE-GPT: Combining Open Access research and large language models for credible, trustworthy question answering

David Pride, Matteo Cancellieri, Petr Knoth

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments 12 pages, accepted submission to TPDL2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11507 2023-06-21 cs.CL cs.AI 81%

TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Yue Huang, Qihui Zhang, Philip S. Y, Lichao Sun

专题命中 安全评测 :trustworthy(title);alignment(abstract);分类 cs.CL、cs.AI

Comments We are currently expanding this work and welcome collaborators!

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08728 2023-06-16 cs.LG cs.AI eess.SP 81%

Towards trustworthy seizure onset detection using workflow notes

Khaled Saab, Siyi Tang, Mohamed Taha, Christopher Lee-Messer, Christopher Ré, Daniel Rubin

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07993 2023-06-16 cs.CR cs.AI cs.LG 81%

Trustworthy Artificial Intelligence Framework for Proactive Detection and Risk Explanation of Cyber Attacks in Smart Grid

Md. Shirajum Munir, Sachin Shetty, Danda B. Rawat

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments Submitted for peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18307 2023-05-31 cs.CY cs.AI 81%

Certification Labels for Trustworthy AI: Insights From an Empirical Mixed-Method Study

Nicolas Scharowski, Michaela Benk, Swen J. Kühne, Léane Wettstein, Florian Brühlmann

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14384 2023-05-25 cs.LG cs.AI cs.CR cs.CV 81%

Adversarial Nibbler: A Data-Centric Challenge for Improving the Safety of Text-to-Image Models

Alicia Parrish, Hannah Rose Kirk, Jessica Quaye, Charvi Rastogi, Max Bartolo, Oana Inel, Juan Ciro, Rafael Mosquera, Addison Howard, Will Cukierski, D. Sculley, Vijay Janapa Reddi, Lora Aroyo

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00470 2023-05-04 cs.LG cs.CY cs.GT 81%

Reward Systems for Trustworthy Medical Federated Learning

Konstantin D. Pandl, Florian Leiser, Scott Thiebes, Ali Sunyaev

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13518 2023-03-24 cs.CV cs.AI cs.LG 81%

Three ways to improve feature alignment for open vocabulary detection

Relja Arandjelović, Alex Andonian, Arthur Mensch, Olivier J. Hénaff, Jean-Baptiste Alayrac, Andrew Zisserman

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07778 2023-03-15 cs.LG cs.AI 81%

GANN: Graph Alignment Neural Network for Semi-Supervised Learning

Linxuan Song, Wenxuan Tu, Sihang Zhou, Xinwang Liu, En Zhu

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09190 2023-02-21 cs.LG cs.CY 81%

Function Composition in Trustworthy Machine Learning: Implementation Choices, Insights, and Questions

Manish Nagireddy, Moninder Singh, Samuel C. Hoffman, Evaline Ju, Karthikeyan Natesan Ramamurthy, Kush R. Varshney

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00646 2023-01-16 cs.SE cs.AI cs.LG cs.RO 81%

Reliability Assessment and Safety Arguments for Machine Learning Components in System Assurance

Yi Dong, Wei Huang, Vibhav Bharti, Victoria Cox, Alec Banks, Sen Wang, Xingyu Zhao, Sven Schewe, Xiaowei Huang

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments Preprint Accepted by ACM Transactions on Embedded Computing Systems

Journal ref ACM Transactions on Embedded Computing Systems; 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03589 2023-01-11 eess.IV cs.AI cs.CV cs.LG 81%

Explainable, Physics Aware, Trustworthy AI Paradigm Shift for Synthetic Aperture Radar

Mihai Datcu, Zhongling Huang, Andrei Anghel, Juanping Zhao, Remus Cacoveanu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00951 2023-01-04 cs.CY cs.AI 81%

Digital Engineering Transformation with Trustworthy AI towards Industry 4.0: Emerging Paradigm Shifts

Jingwei Huang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments Accepted Version 23 pages, 9 figures

Journal ref Transactions of the SDPS: Journal of Integrated Design and Process Science, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10045 2022-10-19 cs.CL cs.AI 81%

SafeText: A Benchmark for Exploring Physical Safety in Language Models

Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, William Yang Wang

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.09239 2022-09-21 cs.LG cs.AI cs.CV 81%

Non-Imaging Medical Data Synthesis for Trustworthy AI: A Comprehensive Survey

Xiaodan Xing, Huanjun Wu, Lichao Wang, Iain Stenson, May Yong, Javier Del Ser, Simon Walsh, Guang Yang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 35 pages, Submitted to ACM Computing Surveys

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14660 2022-09-01 cs.LG cs.AI cs.CV cs.RO cs.SE 81%

Unifying Evaluation of Machine Learning Safety Monitors

Joris Guerin, Raul Sena Ferreira, Kevin Delmas, Jérémie Guiochet

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 9 pages, 5 figures, 3 tables, to appear in the proceedings of the 33rd IEEE International Symposium on Software Reliability Engineering (ISSRE 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.00898 2022-08-02 cs.LG cs.AI cs.CV 81%

Joint covariate-alignment and concept-alignment: a framework for domain generalization

Thuan Nguyen, Boyang Lyu, Prakash Ishwar, Matthias Scheutz, Shuchin Aeron

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 8 pages, 2 figures, and 1 table. This paper is accepted at 32nd IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02009 2022-07-06 cs.LG cs.AI eess.SP 81%

Towards trustworthy Energy Disaggregation: A review of challenges, methods and perspectives for Non-Intrusive Load Monitoring

Maria Kaselimi, Eftychios Protopapadakis, Athanasios Voulodimos, Nikolaos Doulamis, Anastasios Doulamis

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.11981 2022-06-27 cs.AI cs.CY 81%

Never trust, always verify : a roadmap for Trustworthy AI?

Lionel Nganyewou Tidjon, Foutse Khomh

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.06591 2022-06-22 cs.SI cs.CY cs.LG 81%

An Interpretable Graph-based Mapping of Trustworthy Machine Learning Research

Noemi Derzsy, Subhabrata Majumdar, Rajat Malik

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

Comments Accepted in CompleNet-2021 (oral presentation)

Journal ref In: Teixeira, A.S., Pacheco, D., Oliveira, M., Barbosa, H., Gonçalves, B., Menezes, R. (eds) Complex Networks XII. CompleNet-Live 2021. Springer Proceedings in Complexity. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.13229 2022-06-14 cs.CV cs.AI cs.LG cs.SI 81%

Network-level Safety Metrics for Overall Traffic Safety Assessment: A Case Study

Xiwen Chen, Hao Wang, Abolfazl Razi, Brendan Russo, Jason Pacheco, John Roberts, Jeffrey Wishart, Larry Head, Alonso Granados Baca

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01167 2022-05-27 cs.AI cs.LG 81%

Trustworthy AI: From Principles to Practices

Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, Bowen Zhou

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏