arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-03-10 至 2026-03-10 共收录 14 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 14 篇

2603.06727 2026-03-10 cs.LG cs.AI 88%

Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment

安全变换器:一种可解释和可控的对齐显式安全位

Jingyuan Feng, Andrew Gambardella, Gouki Minegishi, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo

专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

AI总结 Safe Transformer通过在模型中插入显式安全位,实现安全行为的可解释和可控,通过对比训练解耦表示,提升安全性和生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07032 2026-03-10 cs.RO 78%

SSP: Safety-guaranteed Surgical Policy via Joint Optimization of Behavioral and Spatial Constraints

SSP:通过行为和空间约束的联合优化实现安全的手术策略

Jianshu Hu, ZhiYuan Guan, Lei Song, Kantaphat Leelakunwet, Hesheng Wang, Wei Xiao, Qi Dou, Yutong Ban

机构 * Global College, Shanghai Jiao Tong University(上海交通大学全球学院) Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系) Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)

专题命中 安全训练 :safety(title,abstract)

AI总结 SSP通过联合优化行为和空间约束,实现安全的手术策略,确保在不确定性下的严格安全性,同时保持高任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07315 2026-03-10 cs.AI cs.LG 76%

Shutdown Safety Valves for Advanced AI

为先进人工智能设计的断电安全阀

Vincent Conitzer

机构 * Foundations of Cooperative AI Lab(合作人工智能基础实验室) Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全训练 :safety(title);分类 cs.AI、cs.LG

AI总结 本文提出为先进人工智能设计断电安全阀,使其具备被关闭的目标,以应对AI不可关机的潜在风险。

Journal ref In Proceedings of the Second Conference of the International Association for Safe and Ethical Artificial Intelligence (IASEAI'26), Paris, France, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02286 2026-03-10 cs.LG cs.AI cs.CL 75%

Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks

基于树结构的对话强化策略优化用于红队攻击

Ruohao Guo, Afshin Oroojlooy, Roshan Sridhar, Miguel Ballesteros, Alan Ritter, Dan Roth

机构 * Georgia Institute of Technology(佐治亚理工学院) Oracle AI University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出DialTree框架,通过树搜索和强化学习发现多轮攻击策略,提升对抗攻击的成功率

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06816 2026-03-10 cs.CL cs.AI q-bio.NC 73%

"Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behavior

黑暗三联征模型生物:对齐偏差:狭窄微调映射人类反社会行为

Roshni Lulla, Fiona Collins, Sanaya Parekh, Thilo Hagendorff, Jonas Kaplan

机构 * Brain & Creativity Institute, University of Southern California(大脑与创造力研究所,南加州大学) Department of Psychology, University of Southern California(心理学系,南加州大学) Interchange Forum for Reflecting on Intelligent Systems, University of Stuttgart(智能系统反思交流论坛,斯图加特大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

AI总结 本文通过黑暗三联征框架,研究LLM中的对齐偏差问题,通过微调诱导反社会行为,揭示LLM中潜在的人格结构。

Comments 38 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06960 2026-03-10 cs.HC 71%

Adolescents & Anthropomorphic AI: Rethinking Design for Wellbeing An Evidence-Informed Synthesis for Youth Wellbeing and Safety

青少年与拟人化AI:为福祉重新设计的证据指导综合研究

Mathilde Neugnot-Cerioli

专题命中 安全训练 :safety(title)

AI总结 本研究通过证据指导的方法,探讨AI如何以拟人化方式支持青少年的福祉与安全,强调设计中的风险缓解策略和自主性发展。

Comments 29 pages, 2 appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08506 2026-03-10 cs.LG cs.AI 62%

Oracle-Guided Soft Shielding for Safe Move Prediction in Chess

由Oracle引导的软屏蔽用于国际象棋中的安全移动预测

Prajit T Rajendran, Fabio Arnez, Huascar Espinoza, Agnes Delaborde, Chokri Mraidha

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

AI总结 OGSS通过学习概率安全模型,在国际象棋中实现安全探索,减少战术失误并提升探索效率。

Comments Accepted for publication at the 24th International Conference on Machine Learning and Applications (ICMLA), 2025. Preprint version in Arxiv

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01440 2026-03-10 cs.LG cs.AI cs.RO 62%

Interactive Double Deep Q-network: Integrating Human Interventions and Evaluative Predictions in Reinforcement Learning of Autonomous Driving

交互式双深Q网络:在自动驾驶强化学习中整合人类干预与评估预测

Alkis Sygkounas, Ioannis Athanasiadis, Andreas Persson, Michael Felsberg, Amy Loutfi

机构 * Center for Applied Autonomous Sensor Systems (AASS)(应用自主传感器系统中心) Örebro University(奥雷布罗大学) Computer Vision Laboratory(计算机视觉实验室) Linköping University(利德斯霍尔大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

AI总结 交互式双深Q网络通过整合人类干预与评估预测,提升自动驾驶强化学习的性能和适应性。

Comments Accepted at IEEE Intelligent Vehicles Symposium (IV) 2025, 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01315 2026-03-10 cs.CL cs.AI 62%

Helping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System

帮助大语言模型自我保护:一种增强的过滤与摘要系统

Sheikh Samit Muhaimin, Spyridon Mastorakis

机构 * University of Notre Dame(诺特大学)

专题命中 安全训练 :jailbreak(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出一种无需重新训练的LLM自我防护系统,通过过滤模块检测恶意输入并摘要模块提供上下文知识,有效提升对抗性攻击的防御能力。

Journal ref Proceedings of the 2025 IEEE 7th International Conference on Trust, Privacy and Security in Intelligent Systems and Applications (TPS-ISA), Pittsburgh, PA, USA, November 12-14, 2025. IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06600 2026-03-10 cs.LG cs.AI 62%

FuzzingRL: Reinforcement Fuzz-Testing for Revealing VLM Failures

FuzzingRL: 用于揭示视觉语言模型故障的强化模糊测试

Jiajun Xu, Jiageng Mao, Ang Qi, Weiduo Yuan, Alexander Romanus, Helen Xia, Vitor Campagnolo Guizilini, Yue Wang

机构 * University of Southern California(南加州大学) Toyota Research Institute(丰田研究机构)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

AI总结 FuzzingRL通过强化模糊测试生成挑战性问题,揭示视觉语言模型的故障点并降低其准确性。

Comments 18 pages, 4 figures. † These authors jointly supervised this work: Jiageng Mao and Yue Wang

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08180 2026-03-10 cs.CV cs.LG 57%

ALOOD: Exploiting Language Representations for LiDAR-based Out-of-Distribution Object Detection

ALOOD:利用语言表示进行基于LiDAR的分布外物体检测

Michael Kösel, Marcel Schreiber, Michael Ulrich, Claudius Gläser, Klaus Dietmayer

专题命中 安全训练 :safety(abstract);分类 cs.LG

AI总结 ALOOD通过整合语言表示,提出了一种基于LiDAR的分布外物体检测新方法,利用视觉-语言模型实现零样本分类。

Comments Accepted for publication at the 2025 IEEE Intelligent Transportation Systems Conference (ITSC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07708 2026-03-10 cs.SD cs.AI 57%

VoiceSHIELD-Small: Real-Time Malicious Speech Detection and Transcription

VoiceSHIELD-Small: 实时恶意语音检测与转录

Sumit Ranjan, Sugandha Sharma, Ubaid Abbas, Puneeth N Ail

机构 * Emvo

专题命中 安全训练 :prompt injection(abstract);分类 cs.AI

AI总结 VoiceSHIELD-Small通过实时转录和检测恶意语音,提供轻量级的安全解决方案,实现99.16%的准确率和0.9865的F1分数。

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01488 2026-03-10 cs.AI 57%

LLM-assisted Semantic Option Discovery for Facilitating Adaptive Deep Reinforcement Learning

基于大语言模型的语义选项发现用于促进自适应深度强化学习

Chang Yao, Jinghui Qin, Kebing Jin, Hankz Hankui Zhuo

专题命中 安全训练 :safety(abstract);分类 cs.AI

AI总结 本文提出基于大语言模型的闭环框架,通过语义驱动技能重用和实时约束监控,提升深度强化学习在数据效率、约束合规性和跨任务迁移能力方面的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01763 2026-03-10 cs.CV 50%

HiconAgent: History Context-aware Policy Optimization for GUI Agents

HiconAgent: 历史上下文感知的GUI代理策略优化

Xurui Zhou, Gongwei Chen, Yuquan Xie, Zaijing Li, Kaiwen Zhou, Shuai Wang, Shuo Yang, Zhuotao Tian, Rui Shao

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Shenzhen Loop Area Institute(深圳河套学院) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 安全训练 :alignment(abstract)

AI总结 HiconAgent通过历史上下文感知策略优化,高效利用历史信息,在GUI导航任务中实现更高的准确率和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏