arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3248 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3248 篇

2510.22514 2026-04-02 eess.SY cs.SY 78%

Robust Multi-Agent Safety via Tube-Based Tightened Exponential Barrier Functions

通过管状紧缩指数屏障函数实现多智能体系统鲁棒安全性

Armel Koulong, Ali Pakniyat

专题命中 安全训练 :safety(title,abstract)

AI总结 本文提出一种构造性框架,用于合成非线性多智能体系统在有界扰动下的可证明安全控制器。核心贡献是通过约束紧缩方法将鲁棒误差反馈与名义轨迹规划相结合,利用RPI管的几何特性推导状态依赖的安全余量,从而保证轨迹安全性。

Comments Joint submission to IFAC World Congress 2026 and NAHS journal (Reference: NAHS_101717). Accepted for NAHS journal; under review by World Congress

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27912 2026-03-31 cs.RO cs.SY eess.SY 78%

Safety Guardrails in the Sky: Realizing Control Barrier Functions on the VISTA F-16 Jet

天空中的安全护栏:在VISTA F-16喷气式飞机上实现控制障碍函数

Andrew W. Singletary, Max H. Cohen, Tamas G. Molnar, Aaron D. Ames

专题命中 安全训练 :safety(title,abstract)

AI总结 本文提出Guardrails机制,通过结合人类或AI指令与安全控制动作,确保自主系统在边缘操作域安全运行,展示了其在F-16飞行测试中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19328 2026-03-23 cs.CR 78%

The Verifier Tax: Horizon Dependent Safety Success Tradeoffs in Tool Using LLM Agents

验证者税:工具使用LLM代理中地平线依赖的安全成功权衡

Tanmay Sah, Vishal Srivastava, Dolly Sah, Kayden Jordan

专题命中 安全训练 :safety(title,abstract)

AI总结 研究运行时强制执行对多步骤工具使用大语言模型代理端到端任务性能的影响,发现安全调解虽能拦截94%的非合规动作,但难以保证安全完成,凸显了需要具备基础身份验证和后干预推理能力的代理。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04904 2026-03-20 cs.HC 78%

Toward Scalable Patient Safety Training: A Prototype for Root Cause Analysis Simulation With AI Virtual Avatars

迈向可扩展的患者安全培训:一种基于AI虚拟仿真的根因分析原型

Yuqi Hu, Qiwen Xiong, Zhenzhen Qin, Brandon Watanabe, Yujing Wang, Mirjana Prpa, Ilmi Yoon

专题命中 安全训练 :safety(title,abstract)

AI总结 本文提出一种AI驱动的仿真平台,通过根因分析模拟提升患者安全培训的可扩展性,利用虚拟角色和AI技术实现沉浸式学习,降低培训成本。

Comments This works has been accepted at the 2026 IEEE Conference on Artificial Intelligence, where a revised version of this work will be published

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16181 2026-03-18 cs.CV cs.CR 78%

KidsNanny: A Two-Stage Multimodal Content Moderation Pipeline Integrating Visual Classification, Object Detection, OCR, and Contextual Reasoning for Child Safety

KidsNanny:一种集成视觉分类、目标检测、OCR和上下文推理的双阶段多模态内容审查流水线,用于儿童安全

Viraj Panchal, Tanmay Talsaniya, Parag Patel, Meet Patel

机构 * KidsNanny Research Team, Vartit Technology Inc.(KidsNanny研究团队,Vartit技术公司)

专题命中 安全训练 :safety(title,abstract)

AI总结 本文提出KidsNanny,一种双阶段多模态内容审查架构,通过视觉Transformer和目标检测器进行视觉筛查,结合OCR和基于文本的7B语言模型进行上下文推理,以提高儿童安全的内容审查效率和准确性。

Comments 12 pages, 2 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06436 2026-03-17 cs.MA 78%

R3R: Decentralized Multi-Agent Collision Avoidance with Infinite-Horizon Safety

R3R:具有无限时间安全性的去中心化多智能体避障

Thomas Marshall Vielmetti, Devansh R. Agrawal, Dimitra Panagou

专题命中 安全训练 :safety(title,abstract)

AI总结 R3R是首个在通信受限条件下提供无限时间安全保证的去中心化多智能体运动规划框架,通过结合安全框架与几何约束实现轨迹的可证明安全。

Comments 8 pages, LaTeX; submitted to the American Control Conference (ACC) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07032 2026-03-10 cs.RO 78%

SSP: Safety-guaranteed Surgical Policy via Joint Optimization of Behavioral and Spatial Constraints

SSP:通过行为和空间约束的联合优化实现安全的手术策略

Jianshu Hu, ZhiYuan Guan, Lei Song, Kantaphat Leelakunwet, Hesheng Wang, Wei Xiao, Qi Dou, Yutong Ban

机构 * Global College, Shanghai Jiao Tong University(上海交通大学全球学院) Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系) Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)

专题命中 安全训练 :safety(title,abstract)

AI总结 SSP通过联合优化行为和空间约束,实现安全的手术策略,确保在不确定性下的严格安全性,同时保持高任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00795 2026-02-25 cs.CV 78%

DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning

DVLA-RL:基于强化学习门控的双层视觉-语言对齐用于少样本学习

Wenhao Li, Xianjing Meng, Qiangchang Wang, Zhongyi Han, Zhibin Wu, Yilong Yin

机构 * Software School, Shandong University(山东大学软件学院) Shenzhen Loop Area Institute(深圳河套学院) School of Computing and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)

专题命中 安全训练 :alignment(title,abstract)

AI总结 DVLA-RL通过双层语义构建和强化学习门控注意力,实现少样本学习中视觉与语言的双层次对齐,提升泛化能力。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13451 2026-02-17 cs.GT 78%

Personalization Aids Pluralistic Alignment Under Competition

个性化有助于竞争中的多元对齐

Natalie Collina, Surbhi Goel, Aaron Roth, Mirah Shi

专题命中 安全训练 :alignment(title,abstract)

AI总结 本文研究了竞争中AI提供者通过个性化策略实现用户多元对齐的可能性,并探讨了不同对齐条件下的均衡结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07629 2026-02-11 cs.RO 78%

LCLA: Language-Conditioned Latent Alignment for Vision-Language Navigation

LCLA:语言引导的潜在对齐用于视觉-语言导航

Nitesh Subedi, Adam Haroon, Samuel Tetteh, Prajwal Koirala, Cody Fleming, Soumik Sarkar

专题命中 安全训练 :alignment(title,abstract)

AI总结 LCLA通过将视觉-语言观测对齐到专家策略的潜在空间,实现轻量级的视觉-运动学习,提升在不同环境下的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07007 2026-02-10 cs.RO 78%

ARGOS: Automated Functional Safety Requirement Synthesis for Embodied AI via Attribute-Guided Combinatorial Reasoning

ARGOS:通过属性引导的组合推理实现具身AI的自动功能安全需求合成

Dongsheng Chen, Yuxuan Li, Yi Lin, Guanhua Chen, Jiaxin Zhang, Xiangyu Zhao, Lei Ma, Xin Yao, Xuetao Wei

机构 * Southern University of Science and Technology, Shenzhen, China(南方科技大学, 深圳, 中国) City University of Hong Kong, Hong Kong, China(香港城市大学, 香港, 中国) The University of Tokyo, Tokyo, Japan(东京大学, 东京, 日本) Lingnan University, Hong Kong, China(岭南大学, 香港, 中国)

专题命中 安全训练 :safety(title,abstract)

AI总结 ARGOS通过属性引导的组合推理,实现具身AI自动功能安全需求生成,提升开放环境中的安全性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12616 2026-01-21 eess.SY cs.SY 78%

Allocating Corrective Control to Mitigate Multi-agent Safety Violations Under Private Preferences

为缓解多智能体安全违规分配纠正控制

Johnathan Corbin, Sarah H. Q. Li, Jonathan Rogers

专题命中 安全训练 :safety(title,abstract)

AI总结 本文提出一种基于隐私保护拍卖机制的多智能体安全纠正控制框架,通过规避信用机制实现高效安全纠正分配。

Comments 8 pages, 3 figures, Submitted to IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20034 2025-12-24 cs.IR 78%

VSA:Visual-Structural Alignment for UI-to-Code

VSA:面向UI到代码的视觉-结构对齐

Xian Wu, Ming Zhang, Zhiyu Fang, Fei Li, Bin Wang, Yong Jiang, Hao Zhou

专题命中 安全训练 :alignment(title,abstract)

AI总结 VSA通过视觉-结构对齐提升UI到代码的模块化和一致性,生成类型安全的组件以提高软件工程效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19238 2025-12-23 cs.CL cs.AI cs.LG 78%

Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation

识别与偏见相关的93个被污名化的群体特征及语言模型的安全防护措施

Anna-Maria Gueorguieva, Aylin Caliskan

专题命中 安全训练 :safety(title);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究探讨了语言模型对93个污名化群体的偏见来源,并测试了防护模型对减少偏见的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10716 2025-12-22 eess.SY cs.SY math.OC 78%

Combinatorial Control Barrier Functions: Nested Boolean and p-choose-r Compositions of Safety Constraints

组合控制障碍函数:安全约束的嵌套布尔和p-choose-r组合

Pio Ong, Haejoon Lee, Tamas G. Molnar, Dimitra Panagou, Aaron D. Ames

专题命中 安全训练 :safety(title,abstract)

AI总结 本文提出组合控制障碍函数框架,用于处理安全约束的嵌套布尔和p-choose-r组合,通过可扩展的方式确保系统安全性。

Comments 6 pages, 3 figures, Accepted for publication in Control System Letters (L-CSS) with the possibility of presenting at the American Control Conference (ACC) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13561 2025-12-16 cs.RO 78%

Near-Field Perception for Safety Enhancement of Autonomous Mobile Robots in Manufacturing Environments

近场感知用于制造环境中自主移动机器人的安全增强

Li-Wei Shih, Ruo-Syuan Mei, Jesse Heidrich, Hui-Ping Wang, Joel Hooton, Joshua Solomon, Jorge Arinez, Guangze Li, Chenhui Shao

机构 * Department of Mechanical Engineering, University of Michigan, MI 48109, USA Materials \& Manufacturing Systems Research, General Motors R \& D, Warren, MI 48092, USA

专题命中 安全训练 :safety(title,abstract)

AI总结 本文提出三级近场感知框架,通过光断续检测、光位移测量和计算机视觉分类,提升自主移动机器人在制造环境中的安全性能。

Comments Submitted to the 54th SME North American Manufacturing Research Conference (NAMRC 54)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13080 2025-12-16 cs.RO 78%

Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos

通过人类视频中的视觉-物理对齐实现空间感知的VLA预训练

Yicheng Feng, Wanpeng Zhang, Ye Wang, Hao Luo, Haoqi Yuan, Sipeng Zheng, Zongqing Lu

机构 * Peking University(北京大学) Renmin University of China(中国人民大学) BeingBeyond

专题命中 安全训练 :alignment(title,abstract)

AI总结 通过人类视频中的视觉-物理对齐实现空间感知的VLA预训练,提升机器人任务的稳健性和通用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15066 2025-12-15 cs.MA cs.IR 78%

Osprey: Production-Ready Agentic AI for Safety-Critical Control Systems

Osprey:面向安全关键控制系统的生产级代理AI

Thorsten Hellert, João Montenegro, Antonin Sulc

专题命中 安全训练 :safety(title,abstract)

AI总结 Osprey通过四种机制解决安全关键控制系统的协调与安全问题,展示在先进光源中的生产应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10506 2025-12-11 cs.SE 78%

SmartC2Rust: Iterative, Feedback-Driven C-to-Rust Translation via Large Language Models for Safety and Equivalence

SmartC2Rust: 基于大语言模型的迭代反馈驱动C到Rust翻译用于安全性和等价性

Momoko Shiraishi, Yinzhi Cao, Takahiro Shinagawa

专题命中 安全训练 :safety(title,abstract)

AI总结 SmartC2Rust通过迭代反馈机制利用大语言模型实现C到Rust的翻译,提升内存安全性和语义等价性。

Journal ref ICSE '26: Proceedings of the 48th International Conference on Software Engineering, April 12-18, 2026, Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05686 2025-12-08 eess.SY cs.SY 78%

LA-RL: Language Action-guided Reinforcement Learning with Safety Guarantees for Autonomous Highway Driving

LA-RL: 基于语言动作引导的安全强化学习用于自动驾驶高速公路驾驶

Yiming Shu, Jiahui Xu, Jiwei Tang, Ruiyang Gao, Chen Sun

专题命中 安全训练 :safety(title,abstract)

AI总结 LA-RL通过整合大语言模型的语义推理和改进的安全层,提升自动驾驶高速公路驾驶的效率与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12610 2025-11-25 eess.SY cs.SY math.OC 78%

Distributed Safe Control Design and Probabilistic Safety Verification for Multi-Agent Systems

分布式安全控制设计与多智能体系统的概率安全验证

Han Wang, Antonis Papachristodoulou, Kostas Margellos

专题命中 安全训练 :safety(title,abstract)

AI总结 本文提出了一种分布式算法,用于多智能体系统的安全控制设计和概率安全验证,通过合作机制解决不可行性问题,并通过场景方法量化安全性。

Comments manuscript accepted by Automatica

Journal ref Automatica, Vol. 179, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14433 2025-11-19 cs.LO cs.RO cs.SE 78%

Safe-ROS: An Architecture for Autonomous Robots in Safety-Critical Domains

Diana C. Benjumea, Marie Farrell, Louise A. Dennis

机构 * Department of Computer Science The University of Manchester Manchester, UK(计算机科学系曼彻斯特大学曼彻斯特英国) University of Manchester Manchester, UK(曼彻斯特大学曼彻斯特英国)

专题命中 安全训练 :safety(title,abstract)

Comments In Proceedings FMAS 2025, arXiv:2511.13245

Journal ref EPTCS 436, 2025, pp. 48-68

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12520 2025-10-28 cs.CV 78%

SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning

Junkai Chen, Zhijie Deng, Kening Zheng, Yibo Yan, Shuliang Liu, PeiJun Wu, Peijie Jiang, Jia Liu, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Southeast University(东南大学) Ant Group, Alibaba(蚂蚁集团)

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02293 2025-10-09 cs.RO cs.MA cs.SY eess.SY 78%

Resolving Conflicting Constraints in Multi-Agent Reinforcement Learning with Layered Safety

Jason J. Choi, Jasmine Jerry Aloor, Jingqi Li, Maria G. Mendoza, Hamsa Balakrishnan, Claire J. Tomlin

机构 * University of California, Berkeley(加州大学伯克利分校) Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全训练 :safety(title,abstract)

Comments Accepted for publication at the 2025 Robotics: Science and Systems Conference. 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02363 2025-10-06 eess.SY cs.SY 78%

Precise HDV Positioning through Safety-Aware Integrated Sensing and Communication in a Value-of-Information-Driven 6G V2X System

Mohammad Reza Abedi, Zahra Rashidi, Nader Mokari, Hamid Saeedi, Nizar Zorba

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24418 2025-09-30 cs.CR 78%

GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners

Haoran Li, Yulin Chen, Jingru Zeng, Hao Peng, Huihao Jing, Wenbin Hu, Xi Yang, Ziqian Zeng, Sirui Han, Yangqiu Song

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13164 2025-09-18 cs.RO cs.SY eess.SY 78%

TeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving

Jiawei Wang, Haowei Sun, Xintao Yan, Shuo Feng, Jun Gao, Henry X. Liu

机构 * University of Michigan(密歇根大学) SaferDrive AI The University of Hong Kong(香港大学) Tsinghua University(清华大学) NVIDIA(英伟达)

专题命中 安全训练 :safety(title,abstract)

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03486 2025-09-12 cs.CR cs.CV cs.SI 78%

UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Yiting Qu, Xinyue Shen, Yixin Wu, Michael Backes, Savvas Zannettou, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA赫尔姆霍茨信息安全中心) TU Delft(代尔夫特理工大学)

专题命中 安全训练 :safety(title,abstract)

Comments To Appear in the ACM Conference on Computer and Communications Security (CCS), October 13, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09423 2025-09-09 cs.RO cs.CV 78%

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Kechun Xu, Xunlong Xia, Kaixuan Wang, Yifei Yang, Yunxuan Mao, Bing Deng, Jieping Ye, Rong Xiong, Yue Wang

机构 * Zhejiang University and Alibaba Cloud(浙江大学和阿里云) Alibaba Cloud(阿里云) Zhejiang University(浙江大学)

专题命中 安全训练 :alignment(title,abstract)

Comments Accepted by T-ASE and CoRL25 GenPriors Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05527 2025-08-08 cs.CV 78%

AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety

Adi Levi, Or Levi, Sardhendu Mishra, Jonathan Morra

机构 * Zefr Inc(Zefr公司)

专题命中 安全训练 :safety(title,abstract)

Comments Accepted to the Computer Vision in Advertising and Marketing (CVAM) workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏