arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1721 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1721 篇

2405.12076 2024-09-09 cs.CR eess.SP 50%

GAN-GRID: A Novel Generative Attack on Smart Grid Stability Prediction

Emad Efatinasab, Alessandro Brighente, Mirco Rampazzo, Nahal Azadi, Mauro Conti

专题命中 越狱攻击 :safety(abstract)

Journal ref European Symposium on Research in Computer Security (ESORICS 2024), 2024, Lecture Notes in Computer Science, vol 14982

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05264 2024-08-13 cs.CR cs.CV 50%

Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security

Yihe Fan, Yuxin Cao, Ziyu Zhao, Ziyao Liu, Shaofeng Li

专题命中 越狱攻击 :trustworthy(abstract)

Comments 8 pages, 1 figure. Accepted to 2024 IEEE International Conference on Systems, Man, and Cybernetics

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13111 2024-07-19 cs.MM cs.CV 50%

PG-Attack: A Precision-Guided Adversarial Attack Framework Against Vision Foundation Models for Autonomous Driving

Jiyuan Fu, Zhaoyu Chen, Kaixun Jiang, Haijing Guo, Shuyong Gao, Wenqiang Zhang

专题命中 越狱攻击 :safety(abstract)

Comments First-Place in the CVPR 2024 Workshop Challenge: Black-box Adversarial Attacks on Vision Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09292 2024-07-18 cs.CR 50%

Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models

Dong Shu, Mingyu Jin, Tianle Chen, Chong Zhang, Yongfeng Zhang

专题命中 越狱攻击 :safety(abstract)

Comments 23 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06948 2024-07-10 eess.SY cs.SY 50%

Detection-Triggered Recursive Impact Mitigation against Secondary False Data Injection Attacks in Microgrids

Mengxiang Liu, Xin Zhang, Rui Zhang, Zhuoran Zhou, Zhenyong Zhang, Ruilong Deng

专题命中 越狱攻击 :trustworthy(abstract)

Comments Submitted to IEEE Transactions on Smart Grid

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02213 2024-07-09 cs.CV 50%

CosPGD: an efficient white-box adversarial attack for pixel-wise prediction tasks

Shashank Agnihotri, Steffen Jung, Margret Keuper

专题命中 越狱攻击 :alignment(abstract)

Comments Accepted at 41st International Conference on Machine Learning (ICML), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19668 2024-05-31 cs.CV 50%

AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization

Jiawei Chen, Xiao Yang, Zhengwei Fang, Yu Tian, Yinpeng Dong, Zhaoxia Yin, Hang Su

专题命中 越狱攻击 :jailbreak(abstract)

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14169 2024-05-24 cs.CV 50%

Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography

Nhat Chung, Sensen Gao, Tuan-Anh Vu, Jie Zhang, Aishan Liu, Yun Lin, Jin Song Dong, Qing Guo

专题命中 越狱攻击 :safety(abstract)

Comments 12 pages, 5 tables, 5 figures, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12916 2024-04-23 cs.CR 50%

Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models

Zhenyang Ni, Rui Ye, Yuxi Wei, Zhen Xiang, Yanfeng Wang, Siheng Chen

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10299 2024-03-26 cs.CV eess.IV 50%

Boosting Adversarial Transferability by Block Shuffle and Rotation

Kunyu Wang, Xuanran He, Wenxuan Wang, Xiaosen Wang

专题命中 越狱攻击 :trustworthy(abstract)

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12693 2024-03-20 cs.CV 50%

As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks?

Anjun Hu, Jindong Gu, Francesco Pinto, Konstantinos Kamnitsas, Philip Torr

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08701 2024-03-20 cs.CR 50%

Review of Generative AI Methods in Cybersecurity

Yagmur Yigit, William J Buchanan, Madjid G Tehrani, Leandros Maglaras

专题命中 越狱攻击 :prompt injection(abstract)

Comments 40 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07654 2024-03-13 cs.IR 50%

Analyzing Adversarial Attacks on Sequence-to-Sequence Relevance Models

Andrew Parry, Maik Fröbe, Sean MacAvaney, Martin Potthast, Matthias Hagen

专题命中 越狱攻击 :prompt injection(abstract)

Comments 13 pages, 3 figures, Accepted at ECIR 2024 as a Full Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15617 2024-02-27 cs.CR cs.SY eess.SY 50%

Reinforcement Learning-Based Approaches for Enhancing Security and Resilience in Smart Control: A Survey on Attack and Defense Methods

Zheyu Zhang

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00994 2024-01-03 cs.CR 50%

Detection and Defense Against Prominent Attacks on Preconditioned LLM-Integrated Virtual Assistants

Chun Fai Chan, Daniel Wankit Yip, Aysan Esmradi

专题命中 越狱攻击 :safety(abstract)

Comments Accepted to be published in the Proceedings of the 10th IEEE CSDE 2023, the Asia-Pacific Conference on Computer Science and Data Engineering 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15736 2023-12-29 cs.CR cs.SY eess.SY 50%

Vulnerability of Machine Learning Approaches Applied in IoT-based Smart Grid: A Review

Zhenyong Zhang, Mengxiang Liu, Mingyang Sun, Ruilong Deng, Peng Cheng, Dusit Niyato, Mo-Yuen Chow, Jiming Chen

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04403 2023-12-08 cs.CV 50%

OT-Attack: Enhancing Adversarial Transferability of Vision-Language Models via Optimal Transport Optimization

Dongchen Han, Xiaojun Jia, Yang Bai, Jindong Gu, Yang Liu, Xiaochun Cao

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18350 2023-12-01 cs.DC 50%

Unveiling Backdoor Risks Brought by Foundation Models in Heterogeneous Federated Learning

Xi Li, Chen Wu, Jiaqi Wang

专题命中 越狱攻击 :safety(abstract)

Comments Jiaqi Wang is the corresponding author. arXiv admin note: text overlap with arXiv:2311.00144

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12685 2023-11-22 cs.CV 50%

Rethinking the Backward Propagation for Adversarial Transferability

Xiaosen Wang, Kangheng Tong, Kun He

专题命中 越狱攻击 :trustworthy(abstract)

Comments Accepted by NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13345 2023-10-23 cs.CR 50%

An LLM can Fool Itself: A Prompt-Based Adversarial Attack

Xilie Xu, Keyi Kong, Ning Liu, Lizhen Cui, Di Wang, Jingfeng Zhang, Mohan Kankanhalli

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10136 2023-09-18 cs.CV 50%

Diversifying the High-level Features for better Adversarial Transferability

Zhiyuan Wang, Zeliang Zhang, Siyuan Liang, Xiaosen Wang

专题命中 越狱攻击 :trustworthy(abstract)

Comments Accepted by BMVC 2023 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15674 2023-08-31 cs.CR 50%

Predict And Prevent DDOS Attacks Using Machine Learning and Statistical Algorithms

Azadeh Golduzian

专题命中 越狱攻击 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11894 2023-08-24 cs.CR cs.CV 50%

Does Physical Adversarial Example Really Matter to Autonomous Driving? Towards System-Level Effect of Adversarial Object Evasion Attack

Ningfei Wang, Yunpeng Luo, Takami Sato, Kaidi Xu, Qi Alfred Chen

专题命中 越狱攻击 :safety(abstract)

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07673 2023-08-16 cs.CV stat.CO stat.ML 50%

A Review of Adversarial Attacks in Computer Vision

Yutong Zhang, Yao Li, Yin Li, Zhichang Guo

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08224 2023-07-18 cs.CR 50%

Uncharted Territory: Energy Attacks in the Battery-less Internet of Things

Luca Mottola, Arslan Hameed, Thiemo Voigt

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09360 2023-07-14 cs.CV 50%

AI Security for Geoscience and Remote Sensing: Challenges and Future Trends

Yonghao Xu, Tao Bai, Weikang Yu, Shizhen Chang, Peter M. Atkinson, Pedram Ghamisi

专题命中 越狱攻击 :safety(abstract)

Journal ref IEEE Geoscience and Remote Sensing Magazine, Volume 11, Issue 2, Pages 60-85, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14782 2023-06-27 cs.CR 50%

On the Resilience of Machine Learning-Based IDS for Automotive Networks

Ivo Zenden, Han Wang, Alfonso Iacovazzi, Arash Vahidi, Rolf Blom, Shahid Raza

专题命中 越狱攻击 :safety(abstract)

Journal ref 2023 IEEE Vehicular Networking Conference (VNC), Istanbul, Turkiye, 2023, pp. 239-246

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.01112 2022-11-29 cs.CR 50%

Adversarial Attack on Radar-based Environment Perception Systems

Amira Guesmi, Ihsen Alouani

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10110 2022-06-22 cs.SE 50%

ProML: A Decentralised Platform for Provenance Management of Machine Learning Software Systems

Nguyen Khoi Tran, Bushra Sabir, M. Ali Babar, Nini Cui, Mehran Abolhasan, Justin Lipman

专题命中 越狱攻击 :safety(abstract)

Comments Accepted as full paper in ECSA 2022 conference. To be presented

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.02675 2021-08-25 cs.CR 50%

Resilience-by-design in Adaptive Multi-Agent Traffic Control Systems

Ranwa Al Mallah, Talal Halabi, Bilal Farooq

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏