arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9311 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9311 篇

2402.13605 2024-06-14 cs.CL 79%

KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge

Jiyoung Lee, Minwoo Kim, Seungho Kim, Junghwan Kim, Seunghyun Won, Hwaran Lee, Edward Choi

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments Accepted at ACL 2024 Findings (35 pages, 7 figures, 16 tables)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07855 2024-06-13 cs.CL cs.SD eess.AS 79%

VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Bing Han, Long Zhou, Shujie Liu, Sanyuan Chen, Lingwei Meng, Yanming Qian, Yanqing Liu, Sheng Zhao, Jinyu Li, Furu Wei

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18510 2024-05-30 cs.AI 79%

Improved Emotional Alignment of AI and Humans: Human Ratings of Emotions Expressed by Stable Diffusion v1, DALL-E 2, and DALL-E 3

James Derek Lomas, Willem van der Maden, Sohhom Bandyopadhyay, Giovanni Lion, Nirmal Patel, Gyanesh Jain, Yanna Litowsky, Haian Xue, Pieter Desmet

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07776 2024-05-29 cs.CL 79%

TELLER: A Trustworthy Framework for Explainable, Generalizable and Controllable Fake News Detection

Hui Liu, Wenya Wang, Haoru Li, Haoliang Li

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

Comments Accepted by Findings of ACL 2024. 28 pages, 2 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12326 2024-05-22 cs.LG 79%

Overlap Number of Balls Model-Agnostic CounterFactuals (ONB-MACF): A Data-Morphology-based Counterfactual Generation Method for Trustworthy Artificial Intelligence

José Daniel Pascual-Triana, Alberto Fernández, Javier Del Ser, Francisco Herrera

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments 21 pages, 6 figures. Submitted to Information Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04937 2024-05-09 cs.AI 79%

Developing trustworthy AI applications with foundation models

Michael Mock, Sebastian Schmidt, Felix Müller, Rebekka Görge, Anna Schmitz, Elena Haedecke, Angelika Voss, Dirk Hecker, Maximillian Poretschkin

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 24 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06676 2024-04-30 cs.LG stat.AP stat.CO 79%

Explainable, Interpretable & Trustworthy AI for Intelligent Digital Twin: Case Study on Remaining Useful Life

Kazuma Kobayashi, Syed Bahauddin Alam

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Journal ref Engineering Applications of Artificial Intelligence 129 (2024): 107620

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14366 2024-04-23 cs.CY 79%

Lessons Learned in Performing a Trustworthy AI and Fundamental Rights Assessment

Marjolein Boonstra, Frédérick Bruneault, Subrata Chakraborty, Tjitske Faber, Alessio Gallucci, Eleanore Hickman, Gerard Kema, Heejin Kim, Jaap Kooiker, Elisabeth Hildt, Annegret Lamadé, Emilie Wiinblad Mathez, Florian Möslein, Genien Pathuis, Giovanni Sartor, Marijke Steege, Alice Stocco, Willy Tadema, Jarno Tuimala, Isabel van Vledder, Dennis Vetter, Jana Vetter, Magnus Westerlund, Roberto V. Zicari

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY

Comments On behalf of the Z-Inspection$^{\small{\circledR}}$ Initiative

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08608 2024-04-15 cs.LG 79%

Hyperbolic Delaunay Geometric Alignment

Aniss Aiman Medbouhi, Giovanni Luca Marchetti, Vladislav Polianskii, Alexander Kravberg, Petra Poklukar, Anastasia Varava, Danica Kragic

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07520 2024-04-15 cs.CV cs.CL 79%

PromptSync: Bridging Domain Gaps in Vision-Language Models through Class-Aware Prototype Alignment and Discrimination

Anant Khandelwal

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments Accepted at CVPR 2024 LIMIT, 12 pages, 8 Tables, 2 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05278 2024-04-15 cs.CL 79%

Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects

Junyu Lu, Dixiang Zhang, Songxin Zhang, Zejian Xie, Zhuoyang Song, Cong Lin, Jiaxing Zhang, Bingyi Jing, Pingjian Zhang

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16908 2024-03-26 cs.AI 79%

Towards Trustworthy Automated Driving through Qualitative Scene Understanding and Explanations

Nassim Belmecheri, Arnaud Gotlieb, Nadjib Lazaar, Helge Spieker

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments SAE International Journal of Connected and Automated Vehicles

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14221 2024-03-25 cs.CL 79%

Improving the Robustness of Large Language Models via Consistency Alignment

Yukun Zhao, Lingyong Yan, Weiwei Sun, Guoliang Xing, Shuaiqiang Wang, Chong Meng, Zhicong Cheng, Zhaochun Ren, Dawei Yin

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments Accepted by LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09565 2024-03-15 cs.SE cs.AI 79%

Welcome Your New AI Teammate: On Safety Analysis by Leashing Large Language Models

Ali Nouri, Beatriz Cabrero-Daniel, Fredrik Törner, Hȧkan Sivencrona, Christian Berger

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments Accepted in CAIN 2024, 6 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05174 2024-03-11 cs.LG 79%

VTruST: Controllable value function based subset selection for Data-Centric Trustworthy AI

Soumi Das, Shubhadip Nag, Shreyyash Sharma, Suparna Bhattacharya, Sourangshu Bhattacharya

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments Accepted in ICLR 2024 DMLR workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03061 2024-03-06 cs.CY 79%

When Industry meets Trustworthy AI: A Systematic Review of AI for Industry 5.0

Eduardo Vyhmeister, Gabriel G. Castane

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY

Comments Review Manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15943 2024-03-05 cs.SE cs.AI 79%

Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware

Ahmed E. Hassan, Dayi Lin, Gopi Krishnan Rajbahadur, Keheliya Gallaba, Filipe R. Cogo, Boyuan Chen, Haoxiang Zhang, Kishanthan Thangarajah, Gustavo Ansaldi Oliva, Jiahuei Lin, Wali Mohammad Abdullah, Zhen Ming Jiang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08208 2024-03-01 cs.AI 79%

Inherent Diverse Redundant Safety Mechanisms for AI-based Software Elements in Automotive Applications

Mandar Pitale, Alireza Abbaspour, Devesh Upadhyay

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments This article is accepted for the SAE WCX 2024 conference proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07547 2024-02-13 cs.MA cs.AI cs.LO cs.SC 79%

Ensuring trustworthy and ethical behaviour in intelligent logical agents

Stefania Costantini

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Journal ref Journal of Logic and Computation, Volume 32, Issue 2, March 2022, Pages 443-478

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02094 2024-02-06 cs.CV cs.AI 79%

Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification

Wenjia Xu, Jiuniu Wang, Zhiwei Wei, Mugen Peng, Yirong Wu

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments Published in ISPRS P&RS. The code is available at https://github.com/wenjiaXu/RS_Scene_ZSL

Journal ref ISPRS Journal of Photogrammetry and Remote Sensing, Volume 198, 2023, Pages 140-152

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09887 2024-02-01 cs.CY 79%

How to Assess Trustworthy AI in Practice

Roberto V. Zicari, Julia Amann, Frédérick Bruneault, Megan Coffee, Boris Düdder, Eleanore Hickman, Alessio Gallucci, Thomas Krendl Gilbert, Thilo Hagendorff, Irmhild van Halem, Elisabeth Hildt, Sune Holm, Georgios Kararigas, Pedro Kringen, Vince I. Madai, Emilie Wiinblad Mathez, Jesmin Jahan Tithi, Dennis Vetter, Magnus Westerlund, Renee Wurth

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY

Comments On behalf of the Z-Inspection$^{\small{\circledR}}$ initiative (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16426 2024-01-31 cs.LG 79%

Informal Safety Guarantees for Simulated Optimizers Through Extrapolation from Partial Simulations

Luke Marks

专题命中 安全评测 :safety(title);alignment(abstract);分类 cs.LG

Comments 17 pages, 0 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15917 2024-01-30 cs.LG cs.CR 79%

Blockchain-enabled Trustworthy Federated Unlearning

Yijing Lin, Zhipeng Gao, Hongyang Du, Jinke Ren, Zhiqiang Xie, Dusit Niyato

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17944 2024-01-26 cs.LG 79%

A Survey on Trustworthy Edge Intelligence: From Security and Reliability To Transparency and Sustainability

Xiaojie Wang, Beibei Wang, Yu Wu, Zhaolong Ning, Song Guo, Fei Richard Yu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments 25 pages, 6 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02062 2024-01-05 stat.ML cs.LG 79%

U-Trustworthy Models.Reliability, Competence, and Confidence in Decision-Making

Ritwik Vashistha, Arya Farahi

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00287 2024-01-02 cs.CL 79%

The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness

Neeraj Varshney, Pavel Dolin, Agastya Seth, Chitta Baral

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10400 2023-12-27 cs.CL cs.CV 79%

What You See is What You Read? Improving Text-Image Alignment Evaluation

Michal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni, Jonathan Herzig, Oran Lang, Eran Ofek, Idan Szpektor

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments Accepted to NeurIPS 2023. Website: https://wysiwyr-itm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09137 2023-12-08 cs.LG stat.ME 79%

Causal prediction models for medication safety monitoring: The diagnosis of vancomycin-induced acute kidney injury

Izak Yasrebi-de Kom, Joanna Klopotowska, Dave Dongelmans, Nicolette De Keizer, Kitty Jager, Ameen Abu-Hanna, Giovanni Cinà

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments Extended Abstract presented at Machine Learning for Health (ML4H) symposium 2023, December 10th, 2023, New Orleans, United States, 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09270 2023-12-01 cs.AI 79%

Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements

Jiawen Deng, Jiale Cheng, Hao Sun, Zhexin Zhang, Minlie Huang

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14859 2023-11-28 cs.LG 79%

An Empirical Investigation into Benchmarking Model Multiplicity for Trustworthy Machine Learning: A Case Study on Image Classification

Prakhar Ganesh

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments Accepted at WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏