arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3266 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3266 篇

2502.15861 2025-02-25 cs.AI cs.CL 73%

C3AI: Crafting and Evaluating Constitutions for Constitutional AI

Yara Kyrychenko, Ke Zhou, Edyta Bogucka, Daniele Quercia

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments This has been accepted for the Web Conference 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08653 2024-12-13 cs.CY cs.AI 73%

What AI evaluations for preventing catastrophic risks can and cannot do

Peter Barnett, Lisa Thiergart

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06843 2024-12-12 cs.CL cs.AI 73%

Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs

Yuxiao Lu, Arunesh Sinha, Pradeep Varakantham

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00380 2024-12-12 cs.CL cs.AI 73%

HonestLLM: Toward an Honest and Helpful Large Language Model

Chujie Gao, Siyuan Wu, Yue Huang, Dongping Chen, Qihui Zhang, Zhengyan Fu, Yao Wan, Lichao Sun, Xiangliang Zhang

专题命中 安全训练 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00033 2024-12-03 cs.AI cs.CY 73%

Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies

Frédéric Berdoz, Roger Wattenhofer

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

Comments Accepted at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02551 2024-10-31 cs.CR cs.AI cs.CY 73%

Breach By A Thousand Leaks: Unsafe Information Leakage in `Safe' AI Responses

David Glukhov, Ziwen Han, Ilia Shumailov, Vardan Papyan, Nicolas Papernot

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09050 2024-09-09 cs.CR cs.AI cs.CV cs.LG 73%

Refusing Safe Prompts for Multi-modal Large Language Models

Zedian Shao, Hongbin Liu, Yuepeng Hu, Neil Zhenqiang Gong

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16254 2024-07-25 cs.CV cs.AI cs.CL cs.MM 73%

Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models

Samuele Poppi, Tobia Poppi, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01198 2024-05-03 cs.LG cs.AI 73%

Towards Interpretable Reinforcement Learning with Constrained Normalizing Flow Policies

Finn Rietz, Erik Schaffernicht, Stefan Heinrich, Johannes A. Stork

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19736 2023-11-28 cs.CL cs.AI 73%

Evaluating Large Language Models: A Comprehensive Survey

Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Supryadi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, Deyi Xiong

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments 111 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18404 2023-07-11 cs.CL cs.LG stat.ML 73%

Conformal Prediction with Large Language Models for Multi-Choice Question Answering

Bhawesh Kumar, Charlie Lu, Gauri Gupta, Anil Palepu, David Bellamy, Ramesh Raskar, Andrew Beam

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.LG

Comments Updated sections on prompt engineering. Expanded sections 4.1 and 4.2 and appendix. Included additional references. Work published at the ICML 2023 (Neural Conversational AI TEACH) workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14553 2023-05-01 cs.AI cs.CL cs.HC 73%

Appropriateness is all you need!

Hendrik Kempt, Alon Lavie, Saskia K. Nagel

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11237 2022-10-21 cs.CR cs.AI cs.LG 73%

Emerging Threats in Deep Learning-Based Autonomous Driving: A Comprehensive Survey

Hui Cao, Wenlong Zou, Yinkun Wang, Ting Song, Mengjun Liu

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments 28 pages,10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.06766 2019-02-20 cs.AI cs.LG stat.ML 73%

Parenting: Safe Reinforcement Learning from Human Input

Christopher Frye, Ilya Feige

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04918 2025-04-08 cs.AI 72%

Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B

Xue Zhang

专题命中 安全训练 :DPO(abstract);harmlessness(abstract);分类 cs.AI;alignment(comments)

Comments 6 pages, 2 figures. Conducted as part of research on alignment techniques for language models

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07686 2026-08-11 quant-ph 新提交 71%

Emergent Problem-Graph Alignment in RL-Discovered Entanglement Topologies for QAOA

用于QAOA的RL发现纠缠拓扑中的涌现问题-图对齐

Tobias Rohe, Federico Harjes Ruiloba, Markus Baumann, Gerhard Stenzel, Leo Sünkel, Thomas Gabor, Claudia Linnhoff-Popien

专题命中 安全训练 :alignment(title)

AI总结 该研究用RL智能体发现QAOA的纠缠拓扑,得到与问题图对齐的稀疏拓扑,在有限优化预算下性能优于全图拓扑,揭示了拓扑密度带来的可训练性与表达性的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06628 2026-07-07 cs.RO cs.CV 版本更新 71%

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment

MIND-V:基于强化学习物理对齐的长期机器人操作分层世界模型

Ruicheng Zhang, Mingyang Zhang, Jun Zhou, Xiaofan Liu, Zunnan Xu, Zhizhou Zhong, Puxin Yan, Haocheng Luo, Xiu Li

机构 * Tsinghua University(清华大学) X Square Robot(X Square机器人) Sun Yat-sen University(中山大学) HKUST(香港科技大学)

专题命中 安全训练 :alignment(title)

AI总结 提出MIND-V分层世界模型,通过语义推理、行为语义桥接和运动视频生成,结合强化学习物理对齐,实现长期机器人操作视频的物理合理合成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07803 2026-06-25 eess.SY cs.SY 新提交 71%

Stable but Unsafe: Agent-Driven Cyber-Physical Systems Under Gain Manipulation Attacks

无安全性的稳定性:对自主网络物理系统的增益操纵攻击

Ali Eslami, Jiangbo Yu

专题命中 安全训练 :safety(title)

AI总结 本文形式化自主网络物理系统中的增益操纵攻击,通过三轴攻击模型和分类法,揭示稳定性保持的增益替换仍可产生远超安全限度的瞬态放大,并利用Bauer-Fike特征值界和Kreiss矩阵定理推导隐蔽条件和最坏影响。

Comments 6 pages, 2 figures, 2 tables. Submitted to IEEE L-CSS for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14351 2026-06-15 cs.CV 新提交 71%

ForceForget: Reinforcement Concept Removal for Enhancing Safety in Text-to-Image Models

ForceForget: 通过强化概念移除增强文本到图像模型的安全性

Dong Han, Yong Li

机构 * Dong Han(董汉) Yong Li(李勇)

专题命中 安全训练 :safety(title)

AI总结 针对文本到图像模型生成不安全内容的问题,提出基于强化学习优化概念擦除奖励的方法,通过安全适配器调节文本嵌入,在消除不安全内容的同时保持模型对安全语义的生成能力。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06940 2026-06-11 eess.AS cs.SD 版本更新 71%

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models

超越语义主导:音频语言模型中的认知情感推理与共情响应对齐

Zhixian Zhao, Shuiyuan Wang, Wenjie Tian, Jingbin Hu, Ziyu Zhang, Lei Xie

机构 * Northwestern Polytechnical University(西北工业大学)

专题命中 安全训练 :alignment(title)

AI总结 提出CogAudio-LLM框架,通过构建LIME-440K数据集实现声学-语义解耦,设计EIPS思维链机制进行心理推理,并采用DR-SAPO优化策略平衡逻辑严谨性与共情质量,解决音频语言模型中的语义主导和情感认知不足问题。

Comments Accepted by Interspeech2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16894 2026-05-19 cs.RO cs.SY eess.SY 71%

Beyond Safety Filtering: Control Barrier Function-Informed Reinforcement Learning for Connected and Automated Vehicles

超越安全过滤:基于控制屏障函数的强化学习用于连接和自动化车辆

Jianye Xu, Bassam Alrifaee

机构 * Department of Computer Science, RWTH Aachen University, Germany(德国亚琛工业大学计算机科学系)

专题命中 安全训练 :safety(title)

AI总结 本文提出了一种基于控制屏障函数的多智能体强化学习奖励设计方法,通过将联合多智能体强化学习动作下的控制屏障函数约束值转化为奖励信号,以显式引导安全学习,并在四向多车道交叉口实验中验证了其在任务性能和对奖励超参数的鲁棒性方面优于传统启发式方法。

Comments This paper has been accepted for publication in the Proceedings of the 2026 IEEE International Conference on Intelligent Transportation Systems (ITSC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06960 2026-03-10 cs.HC 71%

Adolescents & Anthropomorphic AI: Rethinking Design for Wellbeing An Evidence-Informed Synthesis for Youth Wellbeing and Safety

青少年与拟人化AI:为福祉重新设计的证据指导综合研究

Mathilde Neugnot-Cerioli

专题命中 安全训练 :safety(title)

AI总结 本研究通过证据指导的方法,探讨AI如何以拟人化方式支持青少年的福祉与安全,强调设计中的风险缓解策略和自主性发展。

Comments 29 pages, 2 appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14031 2025-11-18 cs.CL cs.AI cs.LG 71%

Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation

Dongyoon Hahm, Taywon Min, Woogyeol Jin, Kimin Lee

专题命中 安全训练 :safety(abstract,comments);分类 cs.CL、cs.AI、cs.LG;alignment(comments)

Comments Accepted at AAAI 2026 AI Alignment Track, Source code: https://github.com/HahmDY/agentic-ft-safety

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24871 2025-10-31 eess.SY cs.SY 71%

Decentralized Merging Control of Connected and Automated Vehicles to Enhance Safety and Energy Efficiency using Control Barrier Functions

Shreshta Rajakumar Deshpande, Mrdjan Jankovic

专题命中 安全训练 :safety(title)

Comments This work has been submitted to a conference for possible publication and is under review. Paper summary: 8 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07651 2025-08-12 eess.SP 71%

Remote ID Based UAV Collision Avoidance Optimization for Low-Altitude Airspace Safety

Ziye Jia, Yian Zhu, Qihui Wu, Lei Zhang, Sen Yang, Zhu Han

专题命中 安全训练 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15371 2025-01-30 cs.RO 71%

Safe and Trustworthy Robot Pathfinding with BIM, MHA*, and NLP

Mani Amani, Reza Akhavian

专题命中 安全训练 :trustworthy(title)

Comments Submitted to IEEE Access

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16511 2024-10-29 cs.CV 71%

GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation

Zhanyu Wang, Longyue Wang, Zhen Zhao, Minghao Wu, Chenyang Lyu, Huayang Li, Deng Cai, Luping Zhou, Shuming Shi, Zhaopeng Tu

专题命中 安全训练 :safety(title)

Comments ACM MM 2024, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19490 2024-03-29 cs.CV 71%

Jointly Training and Pruning CNNs via Learnable Agent Guidance and Alignment

Alireza Ganjdanesh, Shangqian Gao, Heng Huang

专题命中 安全训练 :alignment(title)

Comments IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.04320 2024-01-10 cs.RO 71%

Autonomous robotic re-alignment for face-to-face underwater human-robot interaction

Demetrious T. Kutzke, Ashwin Wariar, Junaed Sattar

专题命中 安全训练 :alignment(title)

Comments Submitted to the Proceedings of the 2024 IEEE Conference on Robotics & Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08638 2023-04-19 cs.MA math.OC 71%

Deep Continuum Deformation Coordination and Optimization with Safety Guarantees

Harshvardhan Uppaluru, Hossein Rastgoftar

专题命中 安全训练 :safety(title)

Comments 6 pages, accepted at ACC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏