arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9311 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9311 篇

2311.12604 2023-11-22 cs.AI 79%

Trustworthy AI: Deciding What to Decide

Caesar Wu, Yuan-Fang Li, Jian Li, Jingjing Xu, Bouvry Pascal

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17594 2023-10-30 cs.CV cs.AI 79%

SPA: A Graph Spectral Alignment Perspective for Domain Adaptation

Zhiqing Xiao, Haobo Wang, Ying Jin, Lei Feng, Gang Chen, Fei Huang, Junbo Zhao

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments NeurIPS 2023 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14492 2023-10-24 cs.CL 79%

Sociocultural Norm Similarities and Differences via Situational Alignment and Explainable Textual Entailment

Sky CH-Wang, Arkadiy Saakyan, Oliver Li, Zhou Yu, Smaranda Muresan

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments EMNLP 2023 Main Conference (Long Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02398 2023-10-05 cs.LG 79%

Reducing Intraspecies and Interspecies Covariate Shift in Traumatic Brain Injury EEG of Humans and Mice Using Transfer Euclidean Alignment

Manoj Vishwanath, Steven Cao, Nikil Dutt, Amir M. Rahmani, Miranda M. Lim, Hung Cao

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05088 2023-09-12 cs.CY q-bio.OT 79%

Towards Trustworthy Artificial Intelligence for Equitable Global Health

Hong Qin, Jude Kong, Wandi Ding, Ramneek Ahluwalia, Christo El Morr, Zeynep Engin, Jake Okechukwu Effoduh, Rebecca Hwa, Serena Jingchuan Guo, Laleh Seyyed-Kalantari, Sylvia Kiwuwa Muyingo, Candace Makeda Moore, Ravi Parikh, Reva Schwartz, Dongxiao Zhu, Xiaoqian Wang, Yiye Zhang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10530 2023-09-08 cs.LG 79%

Reentry Risk and Safety Assessment of Spacecraft Debris Based on Machine Learning

Hu Gao, Zhihui Li, Depeng Dang, Jingfan Yang, Ning Wang

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09217 2023-08-21 cs.CL 79%

Conversational Ontology Alignment with ChatGPT

Sanaz Saki Norouzi, Mohammad Saeid Mahdavinejad, Pascal Hitzler

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16457 2023-08-01 cs.CL 79%

A Benchmark for Understanding Dialogue Safety in Mental Health Support

Huachuan Qiu, Tong Zhao, Anqi Li, Shuai Zhang, Hongliang He, Zhenzhong Lan

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments accepted to The 12th CCF International Conference on Natural Language Processing and Chinese Computing (NLPCC2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08823 2023-07-19 cs.CY 79%

Risk assessment at AGI companies: A review of popular risk assessment techniques from other safety-critical industries

Leonie Koessler, Jonas Schuett

专题命中 安全评测 :safety(title,abstract);分类 cs.CY

Comments 44 pages, 13 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08010 2023-06-05 cs.CL 79%

ProKnow: Process Knowledge for Safety Constrained and Explainable Question Generation for Mental Health Diagnostic Assistance

Kaushik Roy, Manas Gaur, Misagh Soltani, Vipula Rawte, Ashwin Kalyan, Amit Sheth

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Journal ref Front. Big Data, 09 January 2023, Sec. Data Science, Volume 5 - 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17941 2023-05-30 cs.RO cs.AI cs.SY eess.SY 79%

Safety of autonomous vehicles: A survey on Model-based vs. AI-based approaches

Dimia Iberraken, Lounis Adouane

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17873 2023-05-30 cs.HC cs.AI 79%

The Digital Divide in Process Safety: Quantitative Risk Analysis of Human-AI Collaboration

He Wen

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16739 2023-05-29 cs.CL 79%

AlignScore: Evaluating Factual Consistency with a Unified Alignment Function

Yuheng Zha, Yichi Yang, Ruichen Li, Zhiting Hu

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments 19 pages, 5 figures, ACL2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11537 2023-05-22 cs.AI 79%

Trustworthy Federated Learning: A Survey

Asadullah Tariq, Mohamed Adel Serhani, Farag Sallabi, Tariq Qayyum, Ezedin S. Barka, Khaled A. Shuaib

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 45 Pages, 8 Figures, 9 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.11933 2023-04-14 cs.RO cs.AI 79%

Improving safety in physical human-robot collaboration via deep metric learning

Maryam Rezayati, Grammatiki Zanni, Ying Zaoshi, Davide Scaramuzza, Hans Wernher van de Venn

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Journal ref 2022 IEEE 27th International Conference on Emerging Technologies and Factory Automation (ETFA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13151 2023-03-24 cs.AI cs.SE 79%

Defining Quality Requirements for a Trustworthy AI Wildflower Monitoring Platform

Petra Heck, Gerard Schouten

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments Preprint - Paper accepted for CAIN23 - 2nd international conference on AI Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10637 2023-02-22 cs.LG cs.CR 79%

A Survey of Trustworthy Federated Learning with Perspectives on Security, Robustness, and Privacy

Yifei Zhang, Dun Zeng, Jinglong Luo, Zenglin Xu, Irwin King

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.06975 2023-02-15 cs.AI 79%

A Review of the Role of Causality in Developing Trustworthy AI Systems

Niloy Ganguly, Dren Fazlija, Maryam Badar, Marco Fisichella, Sandipan Sikdar, Johanna Schrader, Jonas Wallat, Koustav Rudra, Manolis Koubarakis, Gourab K. Patro, Wadhah Zai El Amri, Wolfgang Nejdl

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 55 pages, 8 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01220 2023-01-06 cs.LG cs.LO cs.SY eess.SY 79%

OVERT: An Algorithm for Safety Verification of Neural Network Control Policies for Nonlinear Systems

Chelsea Sidrane, Amir Maleki, Ahmed Irfan, Mykel J. Kochenderfer

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments 44 pages, under review

Journal ref Journal of Machine Learning Research 23 (2022) 1-45

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.05232 2022-12-20 cs.HC cs.CV cs.LG 79%

Trustworthy Visual Analytics in Clinical Gait Analysis: A Case Study for Patients with Cerebral Palsy

Alexander Rind, Djordje Slijepčević, Matthias Zeppelzauer, Fabian Unglaube, Andreas Kranzl, Brian Horsak

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments 7 pages, 4 figures; supplemental material 9 pages, 8 figures

Journal ref Proceedings of the 2022 IEEE Workshop on TRust and EXpertise in Visual Analytics, TREX (2022) 8-15

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00202 2022-12-07 physics.ao-ph cs.LG 79%

Explainable Artificial Intelligence for Bayesian Neural Networks: Towards trustworthy predictions of ocean dynamics

Mariana C. A. Clare, Maike Sonnewald, Redouane Lguensat, Julie Deshayes, Venkatramani Balaji

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments 25 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02764 2022-12-07 eess.IV cs.CV cs.LG 79%

A Trustworthy Framework for Medical Image Analysis with Deep Learning

Kai Ma, Siyuan He, Pengcheng Xi, Ashkan Ebadi, Stéphane Tremblay, Alexander Wong

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16444 2022-11-30 cs.AI cs.HC 79%

Holding AI to Account: Challenges for the Delivery of Trustworthy AI in Healthcare

Rob Procter, Peter Tolmie, Mark Rouncefield

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 35 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.01267 2022-11-03 cs.CL cs.IR 79%

Multi-Vector Retrieval as Sparse Alignment

Yujie Qian, Jinhyuk Lee, Sai Meher Karthik Duddu, Zhuyun Dai, Siddhartha Brahma, Iftekhar Naim, Tao Lei, Vincent Y. Zhao

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.03848 2022-11-02 stat.ML cs.LG 79%

Time Series Alignment with Global Invariances

Titouan Vayer, Romain Tavenard, Laetitia Chapel, Nicolas Courty, Rémi Flamary, Yann Soullard

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

Comments Published in Transactions on Machine Learning (Oct 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14363 2022-10-27 cs.CL 79%

Enhancing Product Safety in E-Commerce with NLP

Kishaloy Halder, Josip Krapac, Dmitry Goryunov, Anthony Brew, Matti Lyra, Alsida Dizdari, William Gillett, Adrien Renahy, Sinan Tang

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12324 2022-10-25 cs.AI 79%

Trustworthy Human Computation: A Survey

Hisashi Kashima, Satoshi Oyama, Hiromi Arai, Junichiro Mori

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 35 pages, 2 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.07770 2022-10-17 cs.IR cs.AI 79%

Towards Trustworthy AI-Empowered Real-Time Bidding for Online Advertisement Auctioning

Xiaoli Tang, Han Yu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03535 2022-10-10 cs.HC cs.LG 79%

From plane crashes to algorithmic harm: applicability of safety engineering frameworks for responsible ML

Shalaleh Rismani, Renee Shelby, Andrew Smart, Edgar Jatho, Joshua Kroll, AJung Moon, Negar Rostamzadeh

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02260 2022-10-06 cs.MA cs.AI 79%

From Intelligent Agents to Trustworthy Human-Centred Multiagent Systems

Mohammad Divband Soorati, Enrico H. Gerding, Enrico Marchioni, Pavel Naumov, Timothy J. Norman, Sarvapali D. Ramchurn, Bahar Rastegari, Adam Sobey, Sebastian Stein, Danesh Tarpore, Vahid Yazdanpanah, Jie Zhang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments Appears in the Special Issue on Multi-Agent Systems Research in the United Kingdom

Journal ref AI Communications, vol. 35, no. 4, pp. 443-457, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏