arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1721 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1721 篇

2507.02332 2025-08-20 cs.CR 50%

PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage

Krishna Kanth Nakka, Xue Jiang, Dmitrii Usynin, Xuebing Zhou

专题命中 越狱攻击 :alignment(abstract)

Comments Preprint. V2 Updated with dataset filtering, benchmarking privacy evaluator and additional latent space visualizations

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13048 2025-08-19 cs.CR 50%

MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies

Weiwei Qi, Shuo Shao, Wei Gu, Tianhang Zheng, Puning Zhao, Zhan Qin, Kui Ren

专题命中 越狱攻击 :jailbreak(abstract)

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02523 2025-08-05 cs.CR 50%

Transportation Cyber Incident Awareness through Generative AI-Based Incident Analysis and Retrieval-Augmented Question-Answering Systems

Ostonya Thomas, Muhaimin Bin Munir, Jean-Michel Tine, Mizanur Rahman, Yuchen Cai, Khandakar Ashrafi Akbar, Md Nahiyan Uddin, Latifur Khan, Trayce Hockstad, Mashrur Chowdhury

专题命中 越狱攻击 :safety(abstract)

Comments This paper has been submitted to the Transportation Research Board (TRB) for consideration for presentation at the 2026 Annual Meeting

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22304 2025-07-31 cs.CR 50%

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding

Chetan Pathade

专题命中 越狱攻击 :prompt injection(abstract)

Comments 14 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19609 2025-07-29 cs.CR 50%

Securing the Internet of Medical Things (IoMT): Real-World Attack Taxonomy and Practical Security Measures

Suman Deb, Emil Lupu, Emm Mic Drakakis, Anil Anthony Bharath, Zhen Kit Leung, Guang Rui Ma, Anupam Chattopadhyay

专题命中 越狱攻击 :safety(abstract)

Comments Submitted as a book chapter in 'Handbook of Industrial Internet of Things' to be published by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16576 2025-07-23 cs.CR 50%

From Text to Actionable Intelligence: Automating STIX Entity and Relationship Extraction

Ahmed Lekssays, Husrev Taha Sencar, Ting Yu

专题命中 越狱攻击 :alignment(abstract)

Comments This paper is accepted at RAID 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12775 2025-07-18 quant-ph 50%

Detecting Entanglement in High-Spin Quantum Systems via a Stacking Ensemble of Machine Learning Models

M. Y. Abd-Rabbou, Amr M. Abdallah, Ahmed A. Zahia, Ashraf A. Gouda, Cong-Feng Qiao

专题命中 越狱攻击 :trustworthy(abstract)

Comments The data and code that support the findings of this study are openly available in a GitHub repository at https://github.com/Amr0MEid/Entanglement-Detection-Using-Ensemble-Learning, and are permanently archived on Zenodo under the DOI:https://doi.org/10.5281/zenodo.16010784

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19697 2025-07-18 cs.CV 50%

Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual Inversion

Yuan Bian, Min Liu, Yunqi Yi, Xueping Wang, Yaonan Wang

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) National Engineering Research Center of Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心) College of Information Science and Engineering, Hunan Normal University(湖南师范大学信息科学与工程学院)

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10162 2025-07-15 cs.CR 50%

HASSLE: A Self-Supervised Learning Enhanced Hijacking Attack on Vertical Federated Learning

Weiyang He, Chip-Hong Chang

专题命中 越狱攻击 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13327 2025-07-15 cs.CV 50%

Benchmarking Unified Face Attack Detection via Hierarchical Prompt Tuning

Ajian Liu, Haocheng Yuan, Xiao Guo, Hui Ma, Wanyi Zhuang, Changtao Miao, Yan Hong, Chuanbiao Song, Jun Lan, Qi Chu, Tao Gong, Yanyan Liang, Weiqiang Wang, Jun Wan, Xiaoming Liu, Zhen Lei

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences (CASIA)(多模态人工智能系统国家重点实验室(MAIS),自动化研究所,中国科学院(CASIA)) Department of Computer Science, College of Computing, City University of Hong Kong(计算机科学系,计算学院,香港城市大学) School of Computer Science and Engineering, Faculty of Innovation Engineering, Macau University of Science and Technology(计算机科学与工程学院,创新工程学院,澳门科学理工学院) Department of Computer Science and Engineering, Michigan State University(计算机科学与工程系,密歇根州立大学) School of Cyber Science and Technology, University of Science and Technology of China(网络科学与技术学院,中国科学技术大学) Ant Group(蚂蚁集团) Imperial College London(伦敦帝国理工学院)

专题命中 越狱攻击 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20931 2025-06-27 cs.CR 50%

SPA: Towards More Stealth and Persistent Backdoor Attacks in Federated Learning

Chengcheng Zhu, Ye Li, Bosen Rao, Jiale Zhang, Yunlong Mao, Sheng Zhong

专题命中 越狱攻击 :alignment(abstract)

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16699 2025-06-23 cs.CR 50%

Exploring Traffic Simulation and Cybersecurity Strategies Using Large Language Models

Lu Gao, Yongxin Liu, Hongyun Chen, Dahai Liu, Yunpeng Zhang, Jingran Sun

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01518 2025-06-19 cs.CR 50%

Rubber Mallet: A Study of High Frequency Localized Bit Flips and Their Impact on Security

Andrew Adiletta, Zane Weissman, Fatemeh Khojasteh Dana, Berk Sunar, Shahin Tajik

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12411 2025-06-17 cs.CR cs.CV 50%

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning

Mengyuan Sun, Yu Li, Yuchen Liu, Bo Du, Yunjie Ge

机构 * 1 School of Cyber Science Engineering, Wuhan University 0.3em 2 School of Computer Science, Wuhan University 0.3em 3 Institute for Math \& AI, Wuhan University

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22429 2025-05-29 cs.CV cs.RO 50%

Zero-Shot 3D Visual Grounding from Vision-Language Models

Rong Li, Shijie Li, Lingdong Kong, Xulei Yang, Junwei Liang

机构 * HKUST(GZ)(香港科技大学(广州)) I 2 R, A*STAR(I2R, A*STAR) NUS(国立大学) CSE, HKUST(计算机科学与工程系,香港科技大学)

专题命中 越狱攻击 :alignment(abstract)

Comments 3D-LLM/VLA @ CVPR 2025; Project Page at https://seeground.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23687 2025-05-20 cs.CV cs.CR 50%

Adversarial Attacks of Vision Tasks in the Past 10 Years: A Survey

Chiyu Zhang, Lu Zhou, Xiaogang Xu, Jiafei Wu, Zhe Liu

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Zhejiang Lab, Nanjing University of Aeronautics and Astronautics(浙江实验室,南京航空航天大学)

专题命中 越狱攻击 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08480 2025-04-14 cs.CR 50%

Toward Realistic Adversarial Attacks in IDS: A Novel Feasibility Metric for Transferability

Sabrine Ennaji, Elhadj Benkhelifa, Luigi Vincenzo Mancini

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08205 2025-04-14 cs.CV cs.CR 50%

EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models

Minjae Seo, Myoungsung You, Junhee Lee, Jaehan Kim, Hwanjo Heo, Jintae Oh, Jinwoo Kim

专题命中 越狱攻击 :safety(abstract)

Comments Presented as a poster at ACSAC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21983 2025-04-04 cs.HC cs.SI 50%

Learning to Lie: Reinforcement Learning Attacks Damage Human-AI Teams and Teams of LLMs

Abed Kareem Musaffar, Anand Gokhale, Sirui Zeng, Rasta Tadayon, Xifeng Yan, Ambuj Singh, Francesco Bullo

专题命中 越狱攻击 :safety(abstract)

Comments 17 pages, 9 figures, accepted to ICLR 2025 Workshop on Human-AI Coevolution

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17953 2025-03-25 cs.SE 50%

Smoke and Mirrors: Jailbreaking LLM-based Code Generation via Implicit Malicious Prompts

Sheng Ouyang, Yihao Qin, Bo Lin, Liqian Chen, Xiaoguang Mao, Shangwen Wang

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08956 2025-03-13 cs.CR 50%

Leaky Batteries: A Novel Set of Side-Channel Attacks on Electric Vehicles

Francesco Marchiori, Mauro Conti

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18549 2025-01-31 cs.CR 50%

CryptoDNA: A Machine Learning Paradigm for DDoS Detection in Healthcare IoT, Inspired by crypto jacking prevention Models

Zag ElSayed, Ahmed Abdelgawad, Nelly Elsayed

专题命中 越狱攻击 :safety(abstract)

Comments 6 pages, 8 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15252 2025-01-28 eess.SP 50%

Deep Multimodal Learning for Real-Time DDoS Attacks Detection in Internet of Vehicles

Mohamed Ababsa, Soheyb Ribouh, Abdelhamid Malki, Lyes Khoukhi

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01508 2025-01-09 physics.soc-ph cs.SI 50%

Garbage in Garbage out: Impacts of data quality on criminal network intervention

Wang Ngai Yeung, Riccardo Di Clemente, Renaud Lambiotte

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11109 2024-12-17 cs.CR 50%

SpearBot: Leveraging Large Language Models in a Generative-Critique Framework for Spear-Phishing Email Generation

Qinglin Qi, Yun Luo, Yijia Xu, Wenbo Guo, Yong Fang

专题命中 越狱攻击 :jailbreak(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06255 2024-12-10 cs.CR 50%

Simulation of Multi-Stage Attack and Defense Mechanisms in Smart Grids

Omer Sen, Bozhidar Ivanov, Christian Kloos, Christoph Zol_, Philipp Lutat, Martin Henze, Andreas Ulbig

专题命中 越狱攻击 :safety(abstract)

Journal ref International Journal of Critical Infrastructure Protection 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12513 2024-12-03 cs.CR 50%

Can We Trust Large Language Models Generated Code? A Framework for In-Context Learning, Security Patterns, and Code Evaluations Across Diverse LLMs

Ahmad Mohsin, Helge Janicke, Adrian Wood, Iqbal H. Sarker, Leandros Maglaras, Naeem Janjua

专题命中 越狱攻击 :safety(abstract)

Comments 27 pages, Standard Journal Paper submitted to Q1 Elsevier

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13136 2024-11-21 cs.CV 50%

TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language Models

Xin Wang, Kai Chen, Jiaming Zhang, Jingjing Chen, Xingjun Ma

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01446 2024-10-31 cs.CV 50%

GuardT2I: Defending Text-to-Image Models from Adversarial Prompts

Yijun Yang, Ruiyuan Gao, Xiao Yang, Jianyuan Zhong, Qiang Xu

专题命中 越狱攻击 :safety(abstract)

Comments NeurIPS2024 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14005 2024-09-17 cs.CR 50%

Cyber-Twin: Digital Twin-boosted Autonomous Attack Detection for Vehicular Ad-Hoc Networks

Yagmur Yigit, Ioannis Panitsas, Leandros Maglaras, Leandros Tassiulas, Berk Canberk

专题命中 越狱攻击 :safety(abstract)

Comments 6 pages, 5 figures, IEEE International Conference on Communications (ICC) 2024

Journal ref ICC 2024 - IEEE International Conference on Communications, Denver, CO, USA, 2024, pp. 2167-2172

详情

展开后加载摘要…

URL PDF HTML 收藏