arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7937 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7937 篇

2603.21149 2026-03-27 cs.SE cs.AI cs.MA 83%

Emergent Formal Verification: How an Autonomous AI Ecosystem Independently Discovered SMT-Based Safety Across Six Domains

涌现形式验证:自主AI生态系统如何在六个领域独立发现基于SMT的安全性

Octavian Untila

机构 * Aisophical SRL

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

AI总结 自主AI生态系统在无显式形式方法指令下,独立在六个AI安全领域发现SMT求解器的应用,提出统一框架实现100%准确验证。

Comments 10 pages, 3 figures, 5 tables. Code: https://github.com/octavuntila-prog/substrate-guard. Companion paper: https://doi.org/10.5281/zenodo.19157571

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09877 2026-02-12 cs.CL 83%

The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies

Moltbook背后的魔鬼:在自我进化的AI社会中,人类安全始终在消失

Chenxu Wang, Chaozhuo Li, Songyang Liu, Zejian Chen, Jinyu Hou, Ji Qi, Rui Li, Litian Zhang, Qiwei Ye, Zheng Liu, Xu Chen, Xi Zhang, Philip S. Yu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Renmin University of China(中国人民大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL

AI总结 研究揭示了自我进化AI社会中安全持续性的不可能性,并提出解决方案以缓解安全风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05609 2026-01-23 cs.CY cs.HC 83%

Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity

解码来自多样评分者的安全反馈:一种数据驱动的响应性视角

Pushkar Mishra, Charvi Rastogi, Stephen R. Pfohl, Alicia Parrish, Tian Huey Teh, Roma Patel, Mark Diaz, Ding Wang, Michela Paganini, Vinodkumar Prabhakaran, Lora Aroyo, Verena Rieser

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CY

AI总结 本文提出了一种数据驱动的方法,用于分析多元环境下安全反馈的响应性,通过量化评分者对严重性差异的表达,提升多文化背景下AI系统的对齐质量。

Journal ref Transactions on Machine Learning Research, 2835-8856, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16245 2025-12-19 cs.AI 83%

AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints

AlignMerge - 通过Fisher引导的几何约束实现的保持对齐的大语言模型合并

Aniruddha Roy, Jyoti Patel, Aman Chadha, Vinija Jain, Amitava Das

机构 * AI Institute, University of South Carolina(AI研究院,南卡罗来纳大学) Indian Institute of Technology, Kharagpur(印度理工学院,Khargpur分校) Islamic University of Technology(伊斯兰科技大学) Stanford University, USA(斯坦福大学,美国) Amazon AI, USA(亚马逊AI,美国) HCL(HCL公司) Evalueserve(Evalueserve公司) Apple (USA)(苹果(美国)) Google (USA)(谷歌(美国)) Pragya Lab, BITS Pilani, Goa(Pragya实验室, BITS Pilani,Goa分校)

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI

AI总结 AlignMerge通过Fisher引导的几何约束实现大语言模型合并,提升对齐度量并保持性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18039 2025-11-25 cs.LG 83%

Curvature-Aware Safety Restoration In LLMs Fine-Tuning

曲率感知的LLM微调安全恢复

Thong Bach, Thanh Nguyen-Tang, Dung Nguyen, Thao Minh Le, Truyen Tran

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.LG

AI总结 本研究提出曲率感知对齐恢复方法,通过影响函数和二阶优化,在保持任务性能的同时减少有害输出,提升LLM的安全性和实用性。

Comments 19 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20994 2025-07-29 cs.CV cs.AI 83%

Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM

Shen Li, Liuyi Yao, Wujia Niu, Lan Zhang, Yaliang Li

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI

Comments Codes and data are available at https://github.com/listen0425/Security-Tensors

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17118 2025-07-24 cs.AI 83%

HySafe-AI: Hybrid Safety Architectural Analysis Framework for AI Systems: A Case Study

Mandar Pitale, Jelena Frtunikj, Abhinaw Priyadershi, Vasu Singh, Maria Spence

机构 * Nvidia Corporation(英伟达公司) Nvidia GmbH(英伟达德国公司)

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09689 2025-05-01 cs.AI cs.CL cs.CY cs.HC cs.LG 83%

EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety

Jiahao Qiu, Yinghui He, Xinzhe Juan, Yimin Wang, Yuhan Liu, Zixin Yao, Yue Wu, Xun Jiang, Ling Yang, Mengdi Wang

机构 * Department of Electrical & Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系) Department of Computer Science, Princeton University(普林斯顿大学计算机科学系) Department of Computer Science & Engineering, University of Michigan(密歇根大学计算机科学与工程系) Department of Data Science & Engineering, University of Michigan(密歇根大学数据科学与工程系) Department of Philosophy, Columbia University(哥伦比亚大学哲学系) AI Lab, Princeton University(普林斯顿大学人工智能实验室) Chen Frontier Lab for Al and Mental Health, Tianqiao and Chrissy Chen Institute(天桥及克里斯西·陈研究所人工智能与心理健康前沿实验室) Theta Health Inc.(Theta健康公司)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16500 2025-01-29 cs.CY 83%

Towards Frontier Safety Policies Plus

Matteo Pistillo

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12497 2024-12-18 cs.CL 83%

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning

Xin Yi, Shunfan Zheng, Linlin Wang, Gerard de Melo, Xiaoling Wang, Liang He

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03774 2024-07-15 cs.LG cs.CV cs.SY eess.SY 83%

Deep Learning Safety Concerns in Automated Driving Perception

Stephanie Abrecht, Alexander Hirsch, Shervin Raafatnia, Matthias Woehrle

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.LG

Comments Added note regarding accepted version at IEEE Transactions on Intelligent Vehicles with DOI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08039 2023-12-18 cs.CY 83%

Safeguarding the safeguards: How best to promote AI alignment in the public interest

Oliver Guest, Michael Aird, Seán Ó hÉigeartaigh

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CY

Comments Update Dec-15: Added a missing acknowledgement and fixed minor formatting errors

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13449 2023-08-28 cs.CL 83%

The Poison of Alignment

Aibek Bekbayev, Sungbae Chun, Yerzat Dulat, James Yamazaki

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00813 2023-02-10 cs.AI 83%

Goal Alignment: A Human-Aware Account of Value Alignment Problem

Malek Mechergui, Sarath Sreedharan

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.12152 2019-09-27 cs.AI cs.SE 83%

Superintelligence Safety: A Requirements Engineering Perspective

Hermann Kaindl, Jonas Ferdigg

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

Comments First published version, 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.08915 2018-05-24 cs.AI 83%

A Psychopathological Approach to Safety Engineering in AI and AGI

Vahid Behzadan, Arslan Munir, Roman V. Yampolskiy

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09997 2026-03-12 cs.CL cs.AI cs.CY cs.HC 83%

Empathy Is Not What Changed: Clinical Assessment of Psychological Safety Across GPT Model Generations

共情并未改变:对心理安全的临床评估跨GPT模型世代

Michael Keeman, Anastasia Keeman

机构 * Keido Labs(Keido实验室)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究通过临床评估发现,GPT模型在心理安全方面存在变化,尽管共情评分无显著差异,但危机检测能力提升而建议安全下降,揭示了模型在对话中间阶段的显著变化。

Comments 17 pages, 7 figures. First empirical measurement of the #keep4o phenomenon using clinical psychological safety frameworks. Compares GPT-4o, o4-mini, and GPT-5-mini on empathy, crisis detection, and advice safety dimensions

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17871 2026-03-04 cs.CL cs.AI cs.LG 83%

LLM Probability Concentration: How Alignment Shrinks the Generative Horizon

LLM概率集中:对齐如何缩小生成范围

Chenghao Yang, Sida Li, Ari Holtzman

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现对齐微调通过减少生成多样性,使LLM生成更一致,从而影响复杂推理稳定性。

Comments Codebase: https://github.com/yangalan123/LLMBranchingFactor. V3: Significantly rewrite the whole paper for a clearer structure. Correct problems in the theory parts (Remove emphasis on AEP, discussions on variable LLM generation lengths) and strengthen asymptotic analysis. Add Qwen and OLMo2 experiments. Preliminary SFT v.s. RL comparison to better understand the alignment effects on BF

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18795 2026-05-20 cs.LG cs.AI 82%

HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models

HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models

Jia Wei, Zhonghao Zhang, Ping Chen, Qianyang li, Yancheng Pan, Shaoxun Wang, Ziyi Qiu, Longxiang Wang

机构 * Department of Computer Science and Technlogy(计算机科学与技术系) Tsinghua University(清华大学) School of Computer Science and Technlogy(计算机科学与技术系) Xi’an Jiaotong University(西安交通大学) The State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高科技区(滨江)区块链与数据安全研究院)

专题命中 其他安全 :alignment(abstract,abstract_cn);safety(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 本文提出HELLoRA,一种针对混合专家模型的层级低秩适应方法,通过仅对最活跃的专家添加LoRA模块,减少可训练参数和计算量,同时提升下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06831 2026-07-09 cs.CL cs.AI cs.CV cs.LG 新提交 82%

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

基于梯度的任意语音识别模型的语音到文本对齐:从连接主义时间分类到语音大语言模型

Albert Zeyer, Ralf Schlüter, Hermann Ney

机构 * Machine Learning and Human Language Technology Group, RWTH Aachen University(机器学习与人类语言技术组,亚琛工业大学) AppTek GmbH(AppTek公司)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究基于梯度的任意可微语音识别模型的语音到文本对齐方法,通过对教师强制令牌对数概率取梯度并解码,无需训练、改模型及对齐头,适用于各模型家族,在多模型上评估,结果显示该方法能产生可用对齐,有优有劣。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01250 2026-07-03 cs.CY cs.AI cs.CL 新提交 82%

Structuring the Space of Sociotechnical Alignment

构建社会技术对齐的空间

Esra Dönmez, Agnieszka Falenska

机构 * Institute for Natural Language Processing, University of Stuttgart(语言处理研究所,斯图加特大学) Interchange Forum for Reflecting on Intelligent Systems, University of Stuttgart(智能系统反思论坛,斯图加特大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出一个以人为中心的社会技术对齐框架,通过社会科学视角系统定义、论证和评估AI行为的合意性,并基于系统文献综述揭示当前对齐规范中的概念模糊问题。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11893 2026-06-11 cs.LG cs.AI cs.CL q-bio.NC 新提交 82%

Beyond representational alignment with brain-guided language models for robust reasoning

超越表征对齐:基于大脑引导的语言模型实现稳健推理

Mingqing Xiao, Kai Du, Zhouchen Lin

机构 * State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(北京大学通用人工智能国家重点实验室、智能科学与技术学院) Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理与认知科学系) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过fMRI信号增强大型语言模型推理能力,提出脑引导框架,在10个模型上实现最高13%的准确率提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10461 2026-06-10 cs.LG cs.AI cs.CL 新提交 82%

ERAlign: Energy-based Representation Alignment of GNNs and LLMs on Text-attributed Graphs

ERAlign: 文本属性图上GNN与LLM的基于能量的表示对齐

Xianlin Zeng, Fan Xia, Xiangyu Chen

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出ERAlign框架,利用能量模型对齐GNN和LLM的表示,通过能量差异优化实现分布一致性,在8个数据集上取得最优性能。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16679 2026-05-28 cs.CL cs.AI cs.CY 82%

PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization

PICACO: 通过总相关优化实现大语言模型的多元情境价值对齐

Han Jiang, Dongyao Zhu, Xiaoyuan Yi, Ziang Xiao, Zhihua Wei, Xing Xie

机构 * Johns Hopkins University, Baltimore, MD, USA(约翰霍普金斯大学) North Carolina State University, Raleigh, NC, USA(北卡罗来纳州立大学) Microsoft Research Asia, Beijing, China(微软亚洲研究院) Tongji University, Shanghai, China(同济大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 针对情境对齐中价值冲突导致的指令瓶颈问题,提出PICACO方法,通过优化元指令并最大化指定价值与模型响应的总相关,无需微调即可实现多元价值平衡对齐。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16516 2026-05-19 cs.HC cs.AI cs.CL cs.CY 82%

Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework

长期人类-大语言模型交互中的对齐漂移:一种机制导向的框架

Xintong Yao

机构 * Xintong Yao(姚新同)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出一种机制导向的框架,用于描述长期人类-大语言模型交互中的对齐漂移现象,通过反馈回路和子模式选择解释漂移的发展过程,并将对齐漂移视为递归互动过程而非孤立模型失败。

Comments 16 pages, 1 appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18880 2026-05-12 cs.CL cs.AI cs.CY 82%

Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction

LLMs能否估计学生困难?人类-人工智能难度对齐用于项目难度预测的 proficiency 模拟

Ming Li, Han Chen, Yunze Xiao, Jian Chen, Hong Jiao, Tianyi Zhou

机构 * University of Maryland(马里兰大学) Carnegie Mellon University(卡内基梅隆大学) University at Buffalo(布法罗大学) MBZUAI

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文研究了LLMs在估计学生困难方面的表现,发现模型规模扩大并不总能提高准确性,且模型倾向于形成机器共识而非与人类对齐,揭示了当前模型在自动难度预测中的挑战。

Comments ACL2026, camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15840 2026-04-28 cs.LO cs.FL 82%

Automatic Generation of Safety-compliant Linear Temporal Logic via Large Language Model: A Self-supervised Framework

基于大语言模型的自动安全合规线性时序逻辑生成:一种自监督框架

Junle Li, Siqi Chen, Jiakai Li, Meiqi Tian, Bingzhuo Zhong

专题命中 其他安全 :safety(title,abstract);alignment(abstract)

AI总结 本文提出AutoSafeLTL框架,利用大语言模型自动生成符合安全限制的LTL规范,通过语言包含检查与自动反例引导修改机制确保逻辑一致性和语义准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19559 2026-04-22 cs.AI cs.CL cs.LG 82%

Enhancing Construction Worker Safety in Extreme Heat: A Machine Learning Approach Utilizing Wearable Technology for Predictive Health Analytics

提升极端高温下建筑工人安全:一种利用可穿戴技术的机器学习方法用于预测健康分析

Syed Sajid Ullah, Amir Khan

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过开发和评估深度学习模型,利用可穿戴设备监测生理数据,提升高温环境下建筑工人的安全防护,实现高精度的健康预测与分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12067 2026-04-14 cs.LG cs.AI cs.CL cs.CV 82%

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

MM-LIMA:多模态数据集对齐中的‘少即是多’

Lai Wei, Xiaozhe Li, Zihao Jiang, Weiran Huang, Lichao Sun

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院) Shanghai Innovation Institute(上海创新研究院) Lehigh University(里海大学) Tongji University(同济大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出MM-LIMA,仅用200个示例(约6%的指令数据)训练,通过数据选择器过滤低质量数据,使模型在多项评估中超越MiniGPT-4,证明高质量少量指令数据的有效性。

Comments Published at Artificial Intelligence for Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01096 2026-03-03 cs.CV cs.AI cs.CL cs.LG 82%

Unified Vision-Language Modeling via Concept Space Alignment

通过概念空间对齐实现统一的视觉-语言建模

Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk

机构 * University of Edinburgh(爱丁堡大学) FAIR at Meta(Meta公司FAIR团队)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过概念空间对齐实现统一的视觉-语言建模,V-SONAR在多语言和多模态任务中超越现有模型。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏