arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-03-30 至 2026-03-30 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 2 篇

2603.24989 2026-03-30 cs.RO cs.AI 57%

Learning Rollout from Sampling:An R1-Style Tokenized Traffic Simulation Model

从采样中学习:一种R1风格的分词交通仿真模型

Ziyan Wang, Peng Chen, Ding Li, Chiwei Li, Qichao Zhang, Zhongpu Xia, Guizhen Yu

机构 * State Key Laboratory of Intelligent Transportation System, Key Laboratory of Autonomous Transportation Technology for Special Vehicles, Ministry of Industry and Information Technology, School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院,智能交通系统国家重点实验室,特种车辆自主运输技术重点实验室(工业和信息化部)) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所,多模态人工智能系统国家重点实验室)

专题命中 安全训练 :safety(abstract);分类 cs.AI

AI总结 本文提出R1Sim模型,通过分词交通仿真结合熵引导采样和GRPO优化,实现探索与利用的平衡,生成真实安全的多智能体行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10300 2026-03-30 cs.CV 50%

From Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification

从模仿到直觉:用于开放实例视频分类的内在推理

Ke Zhang, Xiangchen Zhao, Yunjie Tian, Jiayu Zheng, Vishal M. Patel, Di Fu

机构 * Johns Hopkins University(约翰霍普金斯大学) ByteDance Inc(字节跳动公司)

专题命中 安全训练 :alignment(abstract)

AI总结 本文提出DeepIntuit框架,通过冷启动监督对齐和GRPO优化,将视频分类从模仿转向直觉推理,提升开放实例分类性能。

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏