arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3248 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3248 篇

2512.13762 2025-12-17 cs.AI cs.HC 79%

State-Dependent Refusal and Learned Incapacity in RLHF-Aligned Language Models

基于状态依赖拒绝和学习无能的强化学习人类反馈对齐语言模型

TK Lee

专题命中 安全训练 :RLHF(title);alignment(abstract);分类 cs.AI

AI总结 本研究提出通过观察行为来审计强化学习人类反馈对齐语言模型中的状态依赖拒绝和学习无能现象。

Comments 23 pages, 6 figures. Qualitative interaction-level analysis of response patterns in a large language model. Code and processed interaction data are available at https://github.com/theMaker-EnvData/llm_learned_incapacity_corpus

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12901 2025-12-16 cs.LG 79%

Predicted-occupancy grids for vehicle safety applications based on autoencoders and the Random Forest algorithm

基于自编码器和随机森林算法的预测占用网格用于车辆安全应用

Parthasarathy Nadarajan, Michael Botsch, Sebastian Sardina

机构 * RMIT University(皇家墨尔本理工大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

AI总结 本文提出基于自编码器和随机森林算法的预测占用网格方法,用于复杂交通场景中的车辆安全应用,通过降维和分类提高预测准确性。

Comments 2017 International Joint Conference on Neural Networks (IJCNN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12387 2025-12-16 cs.LG 79%

Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment

在时间和群体维度上锚定值以实现流匹配模型对齐

Yawen Shao, Jie Xiao, Kai Zhu, Yu Liu, Wei Zhai, Yang Cao, Zheng-Jun Zha

机构 * University of Science and Technology of China(中国科学技术大学) Tongyi Lab(通义实验室)

专题命中 安全训练 :alignment(title,abstract);分类 cs.LG

AI总结 VGPO通过在时间和群体维度上锚定值,改进流匹配模型对齐,提升图像生成质量与任务准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19486 2025-12-12 eess.SY cs.LG cs.MA cs.SY 79%

VLMLight: Safety-Critical Traffic Signal Control via Vision-Language Meta-Control and Dual-Branch Reasoning Architecture

VLMLight:通过视觉-语言元控制与双分支推理架构实现安全关键的交通信号控制

Maonan Wang, Yirong Chen, Aoyu Pang, Yuxin Cai, Chung Shue Chen, Yuheng Kan, Man-On Pun

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shanghai AI Laboratory(上海人工智能实验室) Nanyang Technological University(南洋理工大学) Nokia Bell Labs(诺基亚贝尔实验室) Fourier Intelligence(Fourier智能)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

AI总结 VLMLight通过视觉-语言元控制与双分支推理架构,实现安全关键的交通信号控制,显著减少紧急车辆等待时间并提升系统安全性。

Comments 25 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02682 2025-12-03 cs.MA cs.AI 79%

Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions

超越单体安全:LLM到LLM交互中的风险分类

Piercosma Bisconti, Marcello Galisai, Federico Pierucci, Marcantonio Bracale, Matteo Prandi

机构 * icaro-lab(ICARO实验室)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

AI总结 本文提出从模型级安全向系统级安全的转变,引入ESRH框架,阐述LLM交互中的集体风险并提出InstitutionalAI架构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01247 2025-12-02 cs.CR cs.CY cs.HC 79%

Benchmarking and Understanding Safety Risks in AI Character Platforms

在AI人物平台中进行基准测试与安全风险理解

Yiluo Wei, Peixian Zhang, Gareth Tyson

专题命中 安全训练 :safety(title,abstract);分类 cs.CY

AI总结 该研究首次对AI人物平台进行大规模安全评估,发现其平均不安全响应率高达65.1%,并提出基于机器学习模型的预测方法以提升平台安全性和交互质量。

Comments Accepted to NDSS '26: The Network and Distributed System Security Symposium 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22738 2025-12-01 cs.LG cs.CR 79%

ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning

ShieldAgent: 通过可验证的安全策略推理实现安全代理

Zhaorun Chen, Mintong Kang, Bo Li

机构 * University of Chicago, Chicago IL, USA(芝加哥大学) University of Illinois at Urbana-Champaign, Champaign IL, USA(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

AI总结 ShieldAgent通过逻辑推理实现对代理动作轨迹的安全策略合规性,有效提升代理防护性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17846 2025-11-27 cs.RO cs.AI 79%

Safety Control of Service Robots with LLMs and Embodied Knowledge Graphs

服务机器人中结合大语言模型和具身知识图谱的安全控制

Yong Qi, Gabriel Kyebambo, Siyuan Xie, Wei Shen, Shenghui Wang, Bitao Xie, Bin He, Zhipeng Wang, Shuo Jiang

机构 * School of Electronic Information(电子信息学院) Artificial Intelligence, Shaanxi University of Science(人工智能,陕西科技大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

AI总结 本文提出结合大语言模型与具身知识图谱的方法,提升服务机器人安全控制,通过预定义指令和知识库验证增强安全实践。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14865 2025-11-10 cs.AI cs.FL cs.RO 79%

Joint Verification and Refinement of Language Models for Safety-Constrained Planning

Yunhao Yang, Neel P. Bhatt, William Ward, Zichao Hu, Joydeep Biswas, Ufuk Topcu

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02371 2025-11-05 cs.LG 79%

LUMA-RAG: Lifelong Multimodal Agents with Provably Stable Streaming Alignment

Rohan Wandre, Yash Gajewar, Namrata Patel, Vivek Dhalkari

机构 * Dept. of Computer Engineering(计算机工程系) SIES Graduate School of Technology(SIES技术研究生学院) Bharatiya Vidya Bhavan's Sardar Patel Institute of Technology(巴哈里亚·维达·巴万学院萨达尔·帕特尔技术学院)

专题命中 安全训练 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21254 2025-10-27 cs.AI 79%

Out-of-Distribution Detection for Safety Assurance of AI and Autonomous Systems

Victoria J. Hodge, Colin Paterson, Ibrahim Habli

机构 * Department of Computer Science University of York(计算机科学系约克大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07961 2025-10-15 cs.RO cs.AI 79%

BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models

Peiyan Li, Yixiang Chen, Hongtao Wu, Xiao Ma, Xiangnan Wu, Yan Huang, Liang Wang, Tao Kong, Tieniu Tan

机构 * CASIA(中国科学院自动化研究所) ByteDance Seed(字节跳动种子实验室) UCAS(中国科学院大学) FiveAges NJU(南京大学)

专题命中 安全训练 :alignment(title,abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10823 2025-10-14 cs.AI cs.NE cs.RO 79%

The Irrational Machine: Neurosis and the Limits of Algorithmic Safety

Daniel Howard

机构 * Howard Science Limited, Malvern, UK(霍华德科学有限公司,英国马尔文) QinetiQ Fellow, UK(QinetiQ Fellow,英国) Member of Senior Common Room, Pembroke College, University of Oxford(奥克斯福德大学彭伯里学院高级共同房间成员)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

Comments 41 pages, 17 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05865 2025-10-08 cs.AI cs.CV cs.RO 79%

The Safety Challenge of World Models for Embodied AI Agents: A Review

Lorenzo Baraldi, Zifan Zeng, Chongzhe Zhang, Aradhana Nayak, Hongbo Zhu, Feng Liu, Qunli Zhang, Peng Wang, Shiming Liu, Zheng Hu, Angelo Cangelosi, Lorenzo Baraldi

机构 * University of Pisa(比萨大学) Huawei RAMS Lab(华为RAMS实验室) Technical University of Munich(慕尼黑技术大学) Technical University of Berlin(柏林技术大学) University of Manchester(曼彻斯特大学) University of Modena and Reggio Emilia(莫德纳和雷吉奥艾米利亚大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05156 2025-10-08 cs.SE cs.AI cs.CR 79%

VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation

Lesly Miculicich, Mihir Parmar, Hamid Palangi, Krishnamurthy Dj Dvijotham, Mirko Montanari, Tomas Pfister, Long T. Le

机构 * Google Cloud AI Research(谷歌云人工智能研究) Google DeepMind(谷歌DeepMind)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23614 2025-09-30 cs.AI 79%

PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents

Yaozu Wu, Jizhou Guo, Dongyuan Li, Henry Peng Zou, Wei-Chieh Huang, Yankai Chen, Zhen Wang, Weizhi Zhang, Yangning Li, Meng Zhang, Renhe Jiang, Philip S. Yu

机构 * The University of Tokyo(东京大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校) Zhejiang University(浙江大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15241 2025-09-29 cs.CL 79%

MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety

Yahan Yang, Soham Dan, Shuo Li, Dan Roth, Insup Lee

机构 * University of Pennsylvania(宾夕法尼亚大学) Microsoft(微软公司) Oracle AI

专题命中 安全训练 :safety(title,abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20513 2025-09-26 cs.AI cs.DC 79%

Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems

Samer Alshaer, Ala Khalifeh, Roman Obermaisser

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19688 2025-09-25 cs.RO cs.LG cs.SY eess.SY math.OC 79%

Formal Safety Verification and Refinement for Generative Motion Planners via Certified Local Stabilization

Devesh Nath, Haoran Yin, Glen Chou

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

Comments 10 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18382 2025-09-24 cs.AI 79%

Evaluating the Safety and Skill Reasoning of Large Reasoning Models Under Compute Constraints

Adarsha Balaji, Le Chen, Rajeev Thakur, Franck Cappello, Sandeep Madireddy

机构 * Argonne National Laboratory(阿贡国家实验室) Mathematics and Computer Science Division(数学与计算机科学 division) Data Science and Learning Division(数据科学与学习 division)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16861 2025-09-23 cs.CR cs.AI cs.SE 79%

AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software

Rui Yang, Michael Fu, Chakkrit Tantithamthavorn, Chetan Arora, Gunel Gulmammadova, Joey Chua

机构 * Monash University(墨尔本大学) The University of Melbourne(墨尔本大学) Transurban

专题命中 安全训练 :safety(title);jailbreak(abstract);分类 cs.AI

Comments Accepted to the ASE 2025 International Conference on Automated Software Engineering, Industry Showcase Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06795 2025-09-09 cs.CL 79%

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint

Yanrui Du, Fenglei Fan, Sendong Zhao, Jiawei Cao, Qika Lin, Kai He, Ting Liu, Bing Qin, Mengling Feng

机构 * SCIR Lab, Harbin Institute of Technology, China(哈尔滨工业大学SCIR实验室) City University of Hong Kong(香港城市大学) National University of Singapore(新加坡国立大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12391 2025-09-08 cs.LG 79%

Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL

Junyu Guo, Zhi Zheng, Donghao Ying, Ming Jin, Shangding Gu, Costas Spanos, Javad Lavaei

机构 * University of California Berkeley(加州大学伯克利分校) Virginia Tech(弗吉尼亚理工大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04413 2025-09-05 eess.SY cs.LG cs.MA cs.RO cs.SY math.OC 79%

SAFE--MA--RRT: Multi-Agent Motion Planning with Data-Driven Safety Certificates

Babak Esmaeili, Hamidreza Modares

机构 * Department of Mechanical Engineering, Michigan State University(机械工程系,密歇根州立大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

Comments Submitted to IEEE Transactions on Automation Science and Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16441 2025-08-29 cs.RO cs.AI math.GN 79%

Safe and Efficient Social Navigation through Explainable Safety Regions Based on Topological Features

Victor Toscano-Duran, Sara Narteni, Alberto Carlevaro, Jérôme Guzzi Rocio Gonzalez-Diaz, Maurizio Mongelli

机构 * Department of Applied Mathematics I, University of Seville(应用数学系,塞维利亚大学) CNR-IEIIT Genoa, Italy(意大利热那亚CNR-IEIIT) SUPSI, IDSIA Lugano, Switzerland(瑞士卢加诺SUPSI,IDSIA)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19562 2025-08-28 cs.AI 79%

Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities

Trisanth Srinivasan, Santosh Patapati

机构 * Cyrion Labs(Cyron实验室)

专题命中 安全训练 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02957 2025-08-18 cs.LG cs.SY eess.SY 79%

Embedding Safety into RL: A New Take on Trust Region Methods

Nikola Milosevic, Johannes Müller, Nico Scherf

机构 * Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所) Center for Scalable Data Analytics and Artificial Intelligence(可扩展数据分析与人工智能中心)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

Comments Accepted at ICML 2025

Journal ref Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08926 2025-08-13 cs.AI 79%

Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language Models

Wei Cai, Jian Zhao, Yuchu Jiang, Tianle Zhang, Xuelong Li

机构 * Peking University(北京大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) Northwestern Polytechnical University(西北工业大学) Southeast University(东南大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04377 2025-08-08 cs.CL 79%

PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages

Priyanshu Kumar, Devansh Jain, Akhila Yerukola, Liwei Jiang, Himanshu Beniwal, Thomas Hartvigsen, Maarten Sap

专题命中 安全训练 :safety(title,abstract);分类 cs.CL

Comments Accepted to COLM 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09486 2025-08-01 cs.LG cs.RO 79%

ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning

Yarden As, Bhavya Sukhija, Lenart Treven, Carmelo Sferrazza, Stelian Coros, Andreas Krause

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏