arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7937 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7937 篇

2601.06197 2026-01-13 cs.AI cs.CR 88%

AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation

AI安全措施、生成式AI与潘多拉魔盒:保护企业和个人声誉的AI安全方法

Prasanna Kumar

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI

AI总结 本文提出利用时间一致性学习技术,通过预训练时间卷积网络模型,高效检测生成式AI的黑暗面问题,以提升AI安全性并保护企业和个人声誉。

Comments 10 pages, 3 Figures, 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08842 2025-11-13 cs.AR cs.AI cs.CR 88%

3D Guard-Layer: An Integrated Agentic AI Safety System for Edge Artificial Intelligence

Eren Kurshan, Yuan Xie, Paul Franzon

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI

Comments Resubmitting Re: Arxiv Committee Approval

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23703 2025-07-01 cs.AI 88%

A New Perspective On AI Safety Through Control Theory Methodologies

Lars Ullrich, Walter Zimmer, Ross Greer, Knut Graichen, Alois C. Knoll, Mohan Trivedi

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI

Comments Accepted to be published as part of the 2025 IEEE Open Journal of Intelligent Transportation Systems (OJ-ITS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20471 2025-06-26 cs.CL 88%

Probing AI Safety with Source Code

Ujwal Narayan, Shreyas Chaudhari, Ashwin Kalyan, Tanmay Rajpurohit, Karthik Narasimhan, Ameet Deshpande, Vishvak Murahari

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02205 2025-01-28 cs.SE cs.AI 88%

Swiss Cheese Model for AI Safety: A Taxonomy and Reference Architecture for Multi-Layered Guardrails of Foundation Model Based Agents

Md Shamsujjoha, Qinghua Lu, Dehai Zhao, Liming Zhu

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06369 2024-10-08 cs.CL 88%

Annotation alignment: Comparing LLM and human annotations of conversational safety

Rajiv Movva, Pang Wei Koh, Emma Pierson

专题命中 其他安全 :alignment(title,abstract);safety(title,abstract);分类 cs.CL

Comments EMNLP 2024 (Main). Main text contains 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11314 2024-09-18 cs.CY 88%

The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI Safety

Kristina Fort

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.CY

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.04232 2016-12-07 cs.AI 88%

Review of state-of-the-arts in artificial intelligence with application to AI safety problem

Vladimir Shakirov

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI

Comments version 2 includes grant information

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14668 2026-04-30 cs.DC 88%

A Byzantine Fault Tolerance Approach towards AI Safety

面向AI安全的拜占庭容错方法

John deVadoss, Matthias Artzt

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract)

AI总结 本文借鉴分布式计算中的拜占庭容错机制,提出通过共识机制提升AI系统可靠性与安全性的架构,以应对意外故障和对抗性条件。

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09353 2025-11-18 cs.CR cs.CV 88%

DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt

Yitong Zhang, Jia Li, Liyi Cai, Ge Li

专题命中 其他安全 :alignment(title,abstract);safety(title,abstract)

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20705 2025-09-26 cs.RO 88%

Building Information Models to Robot-Ready Site Digital Twins (BIM2RDT): An Agentic AI Safety-First Framework

Reza Akhavian, Mani Amani, Johannes Mootz, Robert Ashe, Behrad Beheshti

机构 * Department of Civil, Construction, and Environmental Engineering, San Diego State University, San Diego, CA, United States(土木、建设与环境工程系,圣地亚哥州立大学) Department of Electrical and Computer Engineering, University of California, San Diego, San Diego, CA, United States(电气与计算机工程系,加州大学圣地亚哥分校) Department of Mechanical and Aerospace Engineering, University of California, San Diego, San Diego, CA, United States(机械与航空航天工程系,加州大学圣地亚哥分校) Department of Computer Science, San Diego State University, San Diego, CA, United States(计算机科学系,圣地亚哥州立大学)

专题命中 其他安全 :safety(title,abstract);AI safety(title);alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18631 2025-07-28 cs.CR 88%

Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment

Hao Li, Lijun Li, Zhenghao Lu, Xianyi Wei, Rui Li, Jing Shao, Lei Sha

专题命中 其他安全 :alignment(title,abstract);safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14597 2023-09-01 cs.SE 88%

AI Safety Subproblems for Software Engineering Researchers

David Gros, Prem Devanbu, Zhou Yu

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract)

Comments Arxived Apr 2023. Update June 2023 to correct some typos and small text changes. Update Sept 2023, small typos/adjustment, adjust intro to clarify citation analysis focus on HLMI / advanced AI, rerun scripts and tweak handling of unknown venues, add TOSEM, de-anon github and acknowledgements

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15236 2026-04-17 cs.CY cs.AI 87%

Agentic Microphysics: A Manifesto for Generative AI Safety

代理微物理:生成AI安全的宣言

Federico Pierucci, Matteo Prandi, Marcantonio Bracale Syrnikov, Marcello Galisai, Piercosma Bisconti

机构 * DEXAI – Icaro Lab Sant’Anna School of Advanced Studies(DEXAI–伊卡洛实验室圣安娜高级研究学院) DEXAI – Icaro Lab Sapienza University of Rome(DEXAI–伊卡洛实验室罗马萨皮恩扎大学) DEXAI – Icaro Lab VU Amsterdam(DEXAI–伊卡洛实验室阿姆斯特丹伏里克大学) Lucidi DEXAI – Icaro Lab Sapienza University of Rome(Lucidi DEXAI–伊卡洛实验室罗马萨皮恩扎大学)

专题命中 其他安全 :safety(title,abstract);AI safety(title);分类 cs.AI、cs.CY

AI总结 本文提出代理安全研究的新方法,强调交互层面机制对集体风险的影响,通过生成安全方法识别足够机制和设计干预措施。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15886 2025-10-22 cs.CY cs.AI 87%

Combining Cost-Constrained Runtime Monitors for AI Safety

Tim Tian Hua, James Baskerville, Henri Lemoine, Mia Hopman, Aryan Bhatt, Tyler Tracy

机构 * MARS Redwood Research

专题命中 其他安全 :safety(title,abstract);AI safety(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.14235 2022-07-22 cs.LG cs.CY 87%

Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety

Sebastian Houben, Stephanie Abrecht, Maram Akila, Andreas Bär, Felix Brockherde, Patrick Feifel, Tim Fingscheidt, Sujan Sai Gannamaneni, Seyed Eghbal Ghobadi, Ahmed Hammam, Anselm Haselhoff, Felix Hauser, Christian Heinzemann, Marco Hoffmann, Nikhil Kapoor, Falk Kappel, Marvin Klingner, Jan Kronenberger, Fabian Küppers, Jonas Löhdefink, Michael Mlynarski, Michael Mock, Firas Mualla, Svetlana Pavlitskaya, Maximilian Poretschkin, Alexander Pohl, Varun Ravi-Kumar, Julia Rosenzweig, Matthias Rottmann, Stefan Rüping, Timo Sämann, Jan David Schneider, Elena Schulz, Gesina Schwalbe, Joachim Sicking, Toshika Srivastava, Serin Varghese, Michael Weber, Sebastian Wirkert, Tim Wirtz, Matthias Woehrle

专题命中 其他安全 :safety(title,abstract);AI safety(title);分类 cs.CY、cs.LG

Comments 94 pages

Journal ref Fingscheidt, T., Gottschalk, H., Houben, S. (eds) Deep Neural Networks and Data for Automated Driving, Springer, Cham (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.04307 2018-09-11 cs.AI cs.MA 87%

AI Safety and Reproducibility: Establishing Robust Foundations for the Neuropsychology of Human Values

Gopal P. Sarma, Nick J. Hay, Adam Safron

专题命中 其他安全 :safety(title,journal_ref);AI safety(title);alignment(abstract);分类 cs.AI

Comments 5 pages

Journal ref In: Gallina B., Skavhaug A., Schoitsch E., Bitsch F. (eds) Computer Safety, Reliability, and Security. SAFECOMP 2018. Lecture Notes in Computer Science, vol 11094. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04237 2026-04-07 cs.AI cs.CY cs.LG 87%

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems

教育强化学习中的教学安全性:形式化并检测AI辅导系统中的奖励黑客

Oluseyi Olukola, Nick Rahimi

机构 * School of Computing Sciences and Computer Engineering(计算科学与计算机工程学院)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);AI safety(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文提出教育强化学习的教学安全性四层模型及奖励黑客严重性指数,通过模拟实验发现奖励设计不足,需结合约束架构和认知需求以减少奖励黑客问题。

Comments 43 pages, 5 figures. Submitted to the International Journal of Artificial Intelligence in Education (IJAIED)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08761 2026-05-28 stat.ML cs.LG 86%

No Certificate for Alignment: Two Independent Impossibilities and the Pareto Frontier of Achievable Safety Guarantees

对齐无证书:两个独立的不可行性与可实现安全保证的帕累托前沿

Ayushi Agarwal

机构 * Independent Researcher(独立研究者)

专题命中 其他安全 :alignment(title,abstract);safety(title);分类 cs.LG

AI总结 本文通过两个独立的不可行性定理证明,在标准计算复杂性和学习理论假设下,对开放或无界输入域的AI对齐进行形式化认证是不可能的,并刻画了可实现的安全保证的帕累托前沿。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01643 2025-07-08 cs.CY 86%

Third-party compliance reviews for frontier AI safety frameworks

Aidan Homewood, Sophie Williams, Noemi Dreksler, John Lidiard, Malcolm Murray, Lennart Heim, Marta Ziosi, Seán Ó hÉigeartaigh, Michael Chen, Kevin Wei, Christoph Winter, Miles Brundage, Ben Garfinkel, Jonas Schuett

专题命中 其他安全 :safety(title,abstract);AI safety(title);分类 cs.CY

Comments 27 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20968 2025-05-28 cs.CL 86%

Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs

Weixiang Zhao, Yulin Hu, Yang Deng, Jiahe Guo, Xingyu Sui, Xinyang Han, An Zhang, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu

机构 * Harbin Institute of Technology(哈尔滨工业大学) Singapore Management University(新加坡管理大学) National University of Singapore(国立新加坡大学)

专题命中 其他安全 :safety(title,abstract);AI safety(title);分类 cs.CL

Comments To appear at ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15317 2024-04-25 cs.SE cs.HC cs.LG 86%

Concept-Guided LLM Agents for Human-AI Safety Codesign

Florian Geissler, Karsten Roscher, Mario Trapp

专题命中 其他安全 :safety(title,abstract);AI safety(title);分类 cs.LG

Comments 5 pages

Journal ref Proceedings of the AAAI-make Spring Symposium, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10569 2023-07-27 cs.LG cs.AI 86%

Deceptive Alignment Monitoring

Andres Carranza, Dhruv Pai, Rylan Schaeffer, Arnuv Tandon, Sanmi Koyejo

专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments Accepted as BlueSky Oral to 2023 ICML AdvML Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28881 2026-08-06 cs.AI 版本更新 85%

Fragility of Value under Imperfect Alignment

不完美对齐下价值的脆弱性

Winter Cross

专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文针对AI系统与人类价值对齐问题,建立模型明确人类价值函数与代理条件准确度的相关条件,凸显过度优化风险,提出采用quantilizers等限制优化压力的AI设计方案。

Comments 24 pages, 7 figures Added acknowledgement of funding

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16856 2026-05-20 cs.AI 85%

Distributional AGI Safety

分布式AGI安全

Nenad Tomašev, Matija Franklin, Julian Jacobs, Sébastien Krier, Simon Osindero

机构 * Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文提出了一种分布式的AGI安全框架,旨在通过设计和实现虚拟代理沙盒经济来应对群体代理协调带来的安全风险,强调市场机制、可审计性和监管的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13803 2026-04-16 cs.CV cs.AI 85%

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation

Gaslight, Gatekeep, V1-V3: 早期视觉皮层对齐保护视觉语言模型免受阿谀操控

Arya Shah, Vaibhav Tripathi, Mayank Singh, Chaklam Silpasuwanchai

机构 * Indian Institute of Technology Gandhinagar(印度理工学院加尔各答分校) Asian Institute of Technology(亚洲理工学院)

专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文研究了视觉语言模型在高风险场景中的脆弱性,通过评估12种模型的脑部对齐和顺从性,发现早期视觉皮层对齐能有效降低模型对阿谀操控的敏感性。

Comments 28 pages, 9 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14417 2026-03-17 cs.CY cs.AI cs.CL cs.LG 85%

Questionnaire Responses Do not Capture the Safety of AI Agents

问卷回复无法捕捉AI代理的安全性

Max Hellrigel-Holderbaum, Edward James Young

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文指出问卷式评估无法准确衡量AI代理的安全性,因LLM在情景中的表现与实际代理行为存在差异,导致评估方法缺乏有效性。

Comments 31 pages, 11 pages main text

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17846 2025-06-24 cs.AI 85%

Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)

Elija Perrier

机构 * Centre for Quantum Software and Information(量子软件与信息中心) University of Technology, Sydney(悉尼技术大学)

专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

Comments Under review for Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18120 2024-10-03 cs.CL 85%

Exploring Multilingual Concepts of Human Value in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages?

Shaoyang Xu, Weilong Dong, Zishan Guo, Xinwei Wu, Deyi Xiong

专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.CL

Comments EMNLP 2024 findings, code&dataset: https://github.com/shaoyangxu/Multilingual-Human-Value-Concepts

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11137 2023-09-14 cs.AI econ.GN q-fin.EC 85%

Of Models and Tin Men: A Behavioural Economics Study of Principal-Agent Problems in AI Alignment using Large-Language Models

Steve Phelps, Rebecca Ranson

专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

Comments 11 pages, 7 figures. For code see https://github.com/phelps-sg/llm-cooperation Updated with minor corrections: - corrected typo: "mesa-optimiser" instead of "meso-optimiser" - Cited Yang et al (2023) in support of claim that LLMs can solve optimisation problems - Acknowledged Seth Aslin for corrections

详情

展开后加载摘要…

URL PDF HTML 收藏