arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3281 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3281 篇

2511.12160 2025-11-18 cs.RO 50%

Game-Theoretic Safe Multi-Agent Motion Planning with Reachability Analysis for Dynamic and Uncertain Environments (Extended Version)

Wenbin Mai, Minghui Liwang, Xinlei Yi, Xiaoyu Xia, Seyyedali Hosseinalipour, Xianbin Wang

机构 * Department of Electrical and Computer Engineering, National University of Singapore(国立新加坡大学电气与计算机工程系) Department of Control Science and Engineering, Shanghai Institute of Intelligent Science and Technology(上海智能科学与技术研究院控制科学与工程系)

专题命中 安全训练 :safety(abstract)

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10586 2025-11-14 eess.SY cs.RO cs.SY 50%

Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction

Omid Mirzaeedodangeh, Eliot Shekhtman, Nikolai Matni, Lars Lindemann

机构 * Automatic Control Laboratory (IfA)(自动控制实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09813 2025-11-14 cs.HC 50%

I've Seen Enough: Measuring the Toll of Content Moderation on Mental Health

Gabrielle M Gauthier, Eesha Ali, Amna Asim, Sarah Cornell-Maier, Lori A. Zoellner

专题命中 安全训练 :red teaming(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09013 2025-11-13 cs.RO cs.CV 50%

UniMM-V2X: MoE-Enhanced Multi-Level Fusion for End-to-End Cooperative Autonomous Driving

Ziyi Song, Chen Xia, Chenbing Wang, Haibao Yu, Sheng Zhou, Zhisheng Niu

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16240 2025-11-04 cs.RO 50%

Cosmos-Surg-dVRK: World Foundation Model-based Automated Online Evaluation of Surgical Robot Policy Learning

Lukas Zbinden, Nigel Nelson, Juo-Tung Chen, Xinhao Chen, Ji Woong Kim, Mahdi Azizian, Axel Krieger, Sean Huver

机构 * NVIDIA Johns Hopkins University(约翰霍普金斯大学) Stanford University(斯坦福大学)

专题命中 安全训练 :alignment(abstract)

Comments minor metadata and notation fixes; +3 citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23899 2025-10-29 cs.MA cs.RO 50%

Coordinated Autonomous Drones for Human-Centered Fire Evacuation in Partially Observable Urban Environments

Maria G. Mendoza, Addison Kalanther, Daniel Bostwick, Emma Stephan, Chinmay Maheshwari, Shankar Sastry

机构 * Mechanical Engineering University of California, Berkeley(机械工程 加州大学伯克利分校) Computer Sciences University of California, Berkeley(计算机科学 加州大学伯克利分校) Computer Engineering Johns Hopkins University(计算机工程 约翰霍普金斯大学)

专题命中 安全训练 :safety(abstract)

Comments Accepted to IEEE Global Humanitarian Technology Conference (GHTC 2025). 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13939 2025-10-28 cs.CV 50%

Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models

Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, Yuheng Li, Konstantinos Psounis, Xiaofeng Yang

机构 * Department of Computer Science and Informatics, Emory University(计算机科学与信息学系,埃默里大学) Department of Computer Science and Department of Electrical and Computer Engineering, University of Southern California(计算机科学系和电气与计算机工程系,南加州大学) Department of Computer Science, University of Tokyo(计算机科学系,东京大学) Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学) Department of Biomedical Engineering, Georgia Institute of Technology and Emory University(生物医学工程系,佐治亚理工学院和埃默里大学)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15937 2025-10-21 q-fin.RM q-fin.TR 50%

Tail-Safe Stochastic-Control SPX-VIX Hedging: A White-Box Bridge Between AI Sensitivities and Arbitrage-Free Market Dynamics

Jian'an Zhang

专题命中 安全训练 :safety(abstract)

Comments 52 pages; 3 figures; PRIMEarxiv template; fully reproducible artifact (code, configs, plots)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00682 2025-10-21 cs.RO 50%

Immersive Explainability: Visualizing Robot Navigation Decisions through XAI Semantic Scene Projections in Virtual Reality

Jorge de Heuvel, Sebastian Müller, Marlene Wessels, Aftab Akhtar, Christian Bauckhage, Maren Bennewitz

机构 * University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究所) Center for Robotics(机器人中心) University of Mainz(美因茨大学) Fraunhofer Institute for Intelligent Analysis and Information Systems IAIS(弗劳恩霍夫智能分析与信息系统研究所)

专题命中 安全训练 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12861 2025-10-21 cs.RO 50%

Safe Multi-Agent Reinforcement Learning for Behavior-Based Cooperative Navigation

Murad Dawood, Sicong Pan, Nils Dengler, Siqi Zhou, Angela P. Schoellig, Maren Bennewitz

机构 * Humanoid Robots Lab, University of Bonn(波恩大学人形机器人实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02679 2025-10-20 eess.SY cs.SY 50%

A Set-Theoretic Robust Control Approach for Linear Quadratic Games with Unknown Counterparts

Francesco Bianchin, Robert Lefringhausen, Elisa Gaetan, Samuel Tesfazgi, Sandra Hirche

专题命中 安全训练 :safety(abstract)

Comments Accepted for publication in the Proceedings of the 64th IEEE Conference on Decision and Control

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12477 2025-10-15 cs.RO 50%

A Task-Efficient Reinforcement Learning Task-Motion Planner for Safe Human-Robot Cooperation

Gaoyuan Liu, Joris de Winter, Kelly Merckaert, Denis Steckelmacher, Ann Nowe, Bram Vanderborght

机构 * Department of Mechanical Engineering, Vrije Universiteit Brussel(布鲁塞尔自由大学机械工程系) imec Flanders Make(弗拉芒制造) Artificial Intelligence (AI) Lab, Vrije Universiteit Brussel(布鲁塞尔自由大学人工智能实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11185 2025-10-14 cs.HC 50%

Principles of Safe AI Companions for Youth: Parent and Expert Perspectives

Yaman Yu, Mohi, Aishi Debroy, Xin Cao, Karen Rudolph, Yang Wang

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08917 2025-10-13 cs.HC 50%

"I know it's not right, but that's what it said to do": Investigating Trust in AI Chatbots for Cybersecurity Policy

Brandon Lit, Edward Crowder, Daniel Vogel, Hassan Khan

专题命中 安全训练 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07604 2025-10-10 cs.SE 50%

RustAssure: Differential Symbolic Testing for LLM-Transpiled C-to-Rust Code

Yubo Bai, Tapti Palit

专题命中 安全训练 :safety(abstract)

Comments 13 pages to appear in Proceedings of ASE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07206 2025-10-09 cs.CV 50%

EigenScore: OOD Detection using Covariance in Diffusion Models

Shirin Shoushtari, Yi Wang, Xiao Shi, M. Salman Asif, Ulugbek S. Kamilov

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校) University of California, Riverside(加州大学河滨分校)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06666 2025-10-09 math.OC 50%

Trajectory-Optimized Density Control with Flow Matching

Xu Duan, Dongmei Chen

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21292 2025-10-07 cs.SE 50%

Semantic Clustering of Civic Proposals: A Case Study on Brazil's National Participation Platform

Ronivaldo Ferreira, Guilherme da Silva, Carla Rocha, Gustavo Pinto

专题命中 安全训练 :alignment(abstract)

Comments 12 pages, in Portuguese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04076 2025-10-07 cs.RO cs.SY eess.SY 50%

From Shadow to Light: Toward Safe and Efficient Policy Learning Across MPC, DeePC, RL, and LLM Agents

Amin Vahidi-Moghaddam, Sayed Pedram Haeri Boroujeni, Iman Jebellat, Ehsan Jebellat, Niloufar Mehrabi, Zhaojian Li

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00882 2025-10-07 cs.SE 50%

SAFE: Advancing Large Language Models in Leveraging Semantic and Syntactic Relationships for Software Vulnerability Detection

Van Nguyen, Surya Nepal, Tingmin Wu, Xingliang Yuan, Carsten Rudolph

专题命中 安全训练 :safety(abstract)

Journal ref Proceedings of the 20th ACM Asia Conference on Computer and Communications Security (ASIA CCS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01623 2025-10-03 cs.CV cs.RO 50%

VLA-R1: Enhancing Reasoning in Vision-Language-Action Models

Angen Ye, Zeyu Zhang, Boyuan Wang, Xiaofeng Wang, Dapeng Zhang, Zheng Zhu

机构 * GigaAI CASIA Tsinghua University(清华大学)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16144 2025-10-02 eess.SY cs.SY 50%

Safe Event-triggered Gaussian Process Learning for Barrier-Constrained Control

Armin Lederer, Azra Begzadić, Sandra Hirche, Jorge Cortés, Sylvia Herbert

专题命中 安全训练 :safety(abstract)

Comments The first two authors contributed equally to the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26593 2025-10-01 cs.HC 50%

Exploring Large Language Model as an Interactive Sports Coach: Lessons from a Single-Subject Half Marathon Preparation

Kichang Lee

专题命中 安全训练 :safety(abstract)

Comments 23 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21102 2025-09-26 cs.CV 50%

Mammo-CLIP Dissect: A Framework for Analysing Mammography Concepts in Vision-Language Models

Suaiba Amina Salahuddin, Teresa Dorszewski, Marit Almenning Martiniussen, Tone Hovda, Antonio Portaluri, Solveig Thrun, Michael Kampffmeyer, Elisabeth Wetzer, Kristoffer Wickstrøm, Robert Jenssen

机构 * UiT The Arctic University of Norway(乌塔大学极地大学) Technical University of Denmark(技术大学) Østfold Hospital Trust(奥斯fold医院信托) Vestre Viken Hospital Trust(维斯特维肯医院信托) Radboud University Nijmegen Medical Centre(拉德堡德大学奈梅亨医疗中心) The Netherlands Cancer Institute(荷兰癌症研究所) Antoni van Leeuwenhoek Hospital(安东尼·弗莱明医院) University of Copenhagen(哥本哈根大学)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04794 2025-09-22 cs.RO 50%

Runtime Learning of Quadruped Robots in Wild Environments

Yihao Cai, Yanbing Mao, Lui Sha, Hongpeng Cao, Marco Caccamo

机构 * Engineering Technology Division, Wayne State University(韦恩州立大学工程技术系) Department of Computer Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) School of Engineering and Design, Technical University of Munich(慕尼黑技术大学工程与设计学院) School of Electrical Engineering and Computer Science, Washington State University(华盛顿州立大学电气与计算机科学学院)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15154 2025-09-19 cs.CV 50%

MedFact-R1: Towards Factual Medical Reasoning via Pseudo-Label Augmentation

Gengliang Li, Rongyu Chen, Bin Li, Linlin Yang, Guodong Ding

机构 * Baosight(博思特) NUS(新加坡国立大学) SIAT(深圳先进技术研究院) CUC(中国科学技术大学) Microsoft(微软) ANU(澳大利亚国立大学)

专题命中 安全训练 :trustworthy(abstract)

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15099 2025-09-19 eess.SY cs.SY 50%

Digital Twin-based Cooperative Autonomous Driving in Smart Intersections: A Multi-Agent Reinforcement Learning Approach

Taoyuan Yu, Kui Wang, Zongdian Li, Tao Yu, Kei Sakaguchi, Walid Saad

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12378 2025-09-17 eess.SY cs.SY 50%

Platoon-Centric Green Light Optimal Speed Advisory Using Safe Reinforcement Learning

Ruining Yang, Jingyuan Zhou, Qiqing Wang, Jinhao Liang, Kaidi Yang

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12085 2025-09-16 eess.SY cs.SY 50%

Compositional shield synthesis for safe reinforcement learning in partial observability

Steven Carr, Georgios Bakirtzis, Ufuk Topcu

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09953 2025-09-15 cs.RO 50%

Detection of Anomalous Behavior in Robot Systems Based on Machine Learning

Mahfuzul I. Nissan, Sharmin Aktar

机构 * Department of Computer Science University of New Orleans(计算机科学系 新奥尔良大学) Department of Computer Science St Mary's University(计算机科学系 斯坦玛丽大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏