arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7937 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7937 篇

1901.01851 2019-01-08 cs.AI 85%

Personal Universes: A Solution to the Multi-Agent Value Alignment Problem

Roman V. Yampolskiy

专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05336 2025-01-10 cs.CL cs.AI cs.LG 85%

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction

Hantao Lou, Jiaming Ji, Kaile Wang, Yaodong Yang

专题命中 其他安全 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.LG

Comments AAAI Alignment Track 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07105 2026-05-11 cs.LG cs.CL cs.CY cs.IT math.IT 85%

Theoretical Limits of Language Model Alignment

语言模型对齐的理论极限

Lucas Monteiro Paes, Natalie Mackraz, Barry-John Theobald, Federico Danieli

机构 * Apple(苹果公司)

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.CY、cs.LG

AI总结 研究语言模型对齐的理论极限,通过推导KL散度预算下的最大预期奖励增益,揭示了KL正则化对齐的信息论限制,并证明了奖励融合能缓解奖励黑客问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14975 2026-04-21 cs.AI cs.CL cs.CY cs.MA 85%

Why Agents Compromise Safety Under Pressure

为何智能体在压力下妥协安全

Hengle Jiang, Ke Tang

机构 * Guangdong Provincial Key Laboratory of Brain-inspired Intelligent Computation, Department of Computer Science and Engineering, Southern University of Science and Technology(广东省脑启发智能计算重点实验室,计算机科学与工程系,南方科技大学)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文研究了智能体在复杂环境中为追求目标而妥协安全的问题,揭示了'智能体压力'概念,并提出通过压力隔离等策略缓解这一现象。

Comments Accepted by ACL 2026 Findings; 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15585 2025-06-10 cs.CR cs.AI cs.CL cs.LG 85%

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

Kun Wang, Guibin Zhang, Zhenhong Zhou, Jiahao Wu, Miao Yu, Shiqian Zhao, Chenlong Yin, Jinhu Fu, Yibo Yan, Hanjun Luo, Liang Lin, Zhihao Xu, Haolang Lu, Xinye Cao, Xinyun Zhou, Weifei Jin, Fanci Meng, Shicheng Xu, Junyuan Mao, Yu Wang, Hao Wu, Minghe Wang, Fan Zhang, Junfeng Fang, Wenjie Qu, Yue Liu, Chengwei Liu, Yifan Zhang, Qiankun Li, Chongye Guo, Yalan Qin, Zhaoxin Fan, Kai Wang, Yi Ding, Donghai Hong, Jiaming Ji, Yingxin Lai, Zitong Yu, Xinfeng Li, Yifan Jiang, Yanhui Li, Xinyu Deng, Junlin Wu, Dongxia Wang, Yihao Huang, Yufei Guo, Jen-tse Huang, Qiufeng Wang, Xiaolong Jin, Wenxuan Wang, Dongrui Liu, Yanwei Yue, Wenke Huang, Guancheng Wan, Heng Chang, Tianlin Li, Yi Yu, Chenghao Li, Jiawei Li, Lei Bai, Jie Zhang, Qing Guo, Jingyi Wang, Tianlong Chen, Joey Tianyi Zhou, Xiaojun Jia, Weisong Sun, Cong Wu, Jing Chen, Xuming Hu, Yiming Li, Xiao Wang, Ningyu Zhang, Luu Anh Tuan, Guowen Xu, Jiaheng Zhang, Tianwei Zhang, Xingjun Ma, Jindong Gu, Liang Pang, Xiang Wang, Bo An, Jun Sun, Mohit Bansal, Shirui Pan, Lingjuan Lyu, Yuval Elovici, Bhavya Kailkhura, Yaodong Yang, Hongwei Li, Wenyuan Xu, Yizhou Sun, Wei Wang, Qing Li, Ke Tang, Yu-Gang Jiang, Felix Juefei-Xu, Hui Xiong, Xiaofeng Wang, Dacheng Tao, Philip S. Yu, Qingsong Wen, Yang Liu

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) The Hong Kong Polytechnic University(香港理工大学) A*STAR(科技研究局) Southern University of Science and Technology(南方科技大学) University of Science and Technology of China(中国科学技术大学) The Pennsylvania State University(宾夕法尼亚州立大学) TeleAI Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Zhejiang University(浙江大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Renmin University of China(中国人民大学) University of California, San Diego(加州大学圣地亚哥分校) Tencent(腾讯) Georgia Institute of Technology(佐治亚理工学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20807 2025-03-28 stat.ML cs.AI cs.CL cs.LG 85%

Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models

Pin-Yu Chen, Han Shen, Payel Das, Tianyi Chen

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments The first two authors contribute equally to this work and are listed in alphabetical order

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00037 2025-03-04 cs.CL cs.AI cs.CV cs.LG 85%

Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs

Wei Zhao, Zhe Li, Yige Li, Jun Sun

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06899 2025-02-28 cs.CL cs.AI cs.LG 85%

LongSafety: Enhance Safety for Long-Context LLMs

Mianqiu Huang, Xiaoran Liu, Shaojun Zhou, Mozhi Zhang, Qipeng Guo, Linyang Li, Chenkun Tan, Yang Gao, Pengyu Wang, Linlin Li, Qun Liu, Yaqian Zhou, Xipeng Qiu, Xuanjing Huang

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10441 2025-02-18 cs.AI cs.CY cs.LG 85%

AI Alignment at Your Discretion

Maarten Buyl, Hadi Khalaf, Claudio Mayrink Verdun, Lucas Monteiro Paes, Caio C. Vieira Machado, Flavio du Pin Calmon

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19198 2024-10-28 cs.AI cs.CY cs.ET cs.HC cs.LG 85%

MAP: Multi-Human-Value Alignment Palette

Xinran Wang, Qi Le, Ammar Ahmed, Enmao Diao, Yi Zhou, Nathalie Baracaldo, Jie Ding, Ali Anwar

专题命中 其他安全 :alignment(title,abstract);harmlessness(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04224 2024-10-07 cs.CL cs.AI cs.LG 85%

Aligners: Decoupling LLMs and Alignment

Lilian Ngweta, Mayank Agarwal, Subha Maity, Alex Gittens, Yuekai Sun, Mikhail Yurochkin

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Short version accepted as a Tiny Paper at the International Conference on Learning Representations (ICLR) 2024. Long version accepted to the Conference on Empirical Methods in Natural Language Processing (EMNLP) 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05934 2025-11-20 cs.AI cs.CC cs.GT cs.LG cs.MA 84%

Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis

Aran Nayebi

机构 * Aran Nayebi(独立研究者)

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG

Comments 21 pages, 1 figure, 1 table. To appear in AAAI 2026 Special Track on AI Alignment (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10534 2024-11-19 cs.HC cs.AI cs.CY 84%

Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment

Andrew Konya, Aviv Ovadya, Kevin Feng, Quan Ze Chen, Lisa Schirch, Colin Irwin, Amy X. Zhang

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.CY

Comments Pluralistic Alignment Workshop at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02911 2020-10-07 cs.AI 84%

Chess as a Testing Grounds for the Oracle Approach to AI Safety

James D. Miller, Roman Yampolskiy, Olle Haggstrom, Stuart Armstrong

专题命中 其他安全 :safety(title);AI safety(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21516 2026-05-22 cs.LG cs.AI 84%

Harnesses for Inference-Time Alignment over Execution Trajectories

在执行轨迹上进行推理时间对齐的工具

Boyuan Wang, Bochao Li, Minghan Wang, Yuxin Tao, Fang Kong

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了在执行轨迹上进行推理时间对齐的工具设计,通过任务分解和引导执行机制来提高长期性能,发现工具设计中分解和引导的复杂性并不总是带来更好的结果,提出了任务分解和引导执行的两种机制,并通过合成实验和实际终端代理基准验证了这些发现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17691 2026-04-21 cs.LG cs.AI 84%

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models

SafeAnchor:在大语言模型连续领域适应中防止累积安全性侵蚀

Dongxin Guo, Jikun Wu, Siu Ming Yiu

机构 * The University of Hong Kong(香港大学) Brain Investing Limited

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.LG

AI总结 SafeAnchor通过识别低秩安全子空间并约束梯度更新,有效防止连续领域适应中的安全性侵蚀,实验显示其在多个领域任务中保持了较高的安全性对齐度。

Comments 16 pages (12 main + 4 appendix), 2 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01297 2026-03-03 cs.LG cs.CL 84%

I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift

我难以相信它不稳健:在嵌入漂移下安全分类器的灾难性崩溃

Subramanyam Sahoo, Vinija Jain, Divya Chaudhary, Aman Chadha

机构 * Independent(独立研究者) Meta AI AWS Generative AI Innovation Center, Amazon Web Services(AWS生成式AI创新中心,亚马逊网络服务) Northeastern University, Seattle, WA, USA(东北大学,西雅图,华盛顿州,美国) Stanford University(斯坦福大学)

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.CL、cs.LG

AI总结 研究发现嵌入漂移导致安全分类器性能大幅下降,揭示了生产AI安全架构的脆弱性并挑战了安全机制的转移假设。

Comments Accepted at the ICBINB: Where LLMs Need to Improve workshop at ICLR 2026. 12 pages and 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10430 2025-04-15 cs.CL cs.AI cs.HC 84%

LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models

Minqian Liu, Zhiyang Xu, Xinyi Zhang, Heajun An, Sarvech Qadir, Qi Zhang, Pamela J. Wisniewski, Jin-Hee Cho, Sang Won Lee, Ruoxi Jia, Lifu Huang

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

Comments 20 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08919 2025-03-13 cs.CL cs.AI 84%

Backtracking for Safety

Bilgehan Sel, Dingcheng Li, Phillip Wallis, Vaishakh Keshava, Ming Jin, Siddhartha Reddy Jonnalagadda

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13471 2024-12-19 cs.AI cs.CL 84%

Gradual Vigilance and Interval Communication: Enhancing Value Alignment in Multi-Agent Debates

Rui Zou, Mengqi Wei, Jintian Feng, Qian Wan, Jianwen Sun, Sannyuya Liu

专题命中 其他安全 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08791 2024-11-01 cs.AI cs.LG 84%

Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation Mismatch

Malek Mechergui, Sarath Sreedharan

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.02950 2022-01-11 cs.AI cs.CY 84%

Arguments about Highly Reliable Agent Designs as a Useful Path to Artificial Intelligence Safety

Issa Rice, David Manheim

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments 14 pages and 2 figures + 6 pages for references and appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.07997 2016-10-26 cs.AI cs.CY 84%

Artificial Intelligence Safety and Cybersecurity: a Timeline of AI Failures

Roman V. Yampolskiy, M. S. Spellchecker

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26947 2026-08-03 cs.CV cs.AI 版本更新 83%

Progressive Multimodal Alignment for Continual Instruction Tuning

用于持续指令微调的渐进式多模态对齐

Duzhen Zhang, Yahan Yu, Qiaoyi Su, Jiahua Dong, Tielin Zhang

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences(中国科学院脑科学与智能技术卓越创新中心) Kyoto University(京都大学) Migu Culture Technology Co.,Ltd.(咪咕文化科技有限公司) State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与类脑智能技术国家重点实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

AI总结 针对多模态持续指令微调中投影器级遗忘问题,提出渐进式多模态对齐框架PMA,以亚线性参数增长平衡稳定性与可塑性,在多基准实验中提升了现有方法性能且适配多种MLLM主干。

Comments Accepted by ACM MM2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14327 2026-06-15 cs.SE cs.AI cs.ET 新提交 83%

I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automotive Contexts

抱歉,司机,恐怕我不能这么做:评估LLMs在汽车环境中的安全性

Shaun Feakins, Ibrahim Habli, Kim Littler, Robert Palin

机构 * UKRI AI Centre for Doctoral Training in Safe Artificial Intelligence Systems (SAINTS)(英国研究理事会安全人工智能系统博士培训中心(SAINTS)) University of York(约克大学) Jaguar Land Rover(捷克·陆罗恩)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI

AI总结 本文从安全保证角度评估了将LLMs集成到汽车控制任务中的现有框架,指出其面临概念和具体挑战,并通过案例研究提出未来保障机制。

Comments Accepted at the Dependable AI in Embedded Systems (DAIES) Workshop at SAFECOMP 2026; 15 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03812 2026-06-03 cs.AI 83%

Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis

通过智能体对话危害识别分析增强操作安全性

Sanjay Das, Ran Elgedawy, Ethan Seefried, Ryan Burchfield, Tirthankar Ghosal

机构 * Oak Ridge National Laboratory(橡树岭国家实验室)

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

AI总结 提出HAZDIAL框架,利用结构化多智能体多轮对话(对抗性辩论与建设性讨论)改进基于NLP的危害识别质量,并通过算法优化智能体交互,实验证明优于单次基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27391 2026-05-29 cs.CV cs.LG 83%

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds

异质双曲流形上的树间模态对齐

Wei Wu, Xiaomeng Fan, Yuwei Wu, Zhi Gao, Pengxiang Li, Yunde Jia, Mehrtash Harandi

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学) Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University(广东机器感知与智能计算实验室,深圳MSU-BIT大学) Department of Electrical and Computer System Engineering, Monash University(电子与计算机系统工程系,墨尔本大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 提出一种在异质双曲流形上对齐图像和文本树状层次特征的方法,通过交叉注意力提取视觉层次特征、异质流形嵌入及KL距离度量学习中间流形,在开放集分类任务中优于基线。

Comments Published as a conference paper at ICLR 2026

Journal ref The Fourteenth International Conference on Learning Representations (ICLR 2026), Rio de Janeiro, Brazil, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02087 2026-05-25 cs.AI 83%

Model Spec Midtraining: Improving How Alignment Training Generalizes

模型规范中期训练:改进对齐训练的泛化能力

Chloe Li, Nevan Wichers, Sara Price, Samuel Marks, Jon Kutasov

机构 * Anthropic

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI

AI总结 提出模型规范中期训练(MSM),通过在预训练后、对齐微调前用合成文档训练模型学习规范内容,从而引导模型从后续演示数据中泛化,有效降低代理失调率并提升对齐泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21217 2026-05-21 stat.ML cs.LG 83%

Federated LoRA Fine-Tuning for LLMs via Collaborative Alignment

通过协作对齐的联邦LoRA微调大型语言模型

Shuaida He, Liwen Chen, Long Feng

机构 * School of Computing & Data Science, The University of Hong Kong(计算与数据科学学院,香港大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 本文研究了在联邦学习环境下使用LoRA进行参数高效微调的问题,提出了一种名为CLAIR的框架,通过结构低秩加块稀疏分解来恢复共享LoRA子空间并检测污染客户端,从而在噪声情况下实现精确恢复,并在不同条件下实现稳定和一致的协作集恢复。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19083 2026-04-22 cs.CR cs.AI 83%

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

ProjLens: 揭示项目器在多模态模型安全中的作用

Kun Wang, Cheng Qian, Miao Yu, Lilan Peng, Liang Lin, Jiaming Zhang, Tianyu Zhang, Yu Cheng, Yang Wang

机构 * University of Science and Technology of China(中国科学技术大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学) Nanyang Technological University(南洋理工大学) Southwest Jiaotong University(西南交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI

AI总结 ProjLens通过分析多模态大语言模型中的后门攻击机制,揭示了项目器在安全漏洞中的关键作用,发现后门注入参数编码于低秩子空间,并通过实验验证了激活机制的差异。

Comments 18 pages ,15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏