arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3281 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3281 篇

2505.06620 2025-05-13 cs.HC cs.AI 57%

Integrating Explainable AI in Medical Devices: Technical, Clinical and Regulatory Insights and Recommendations

Dima Alattal, Asal Khoshravan Azar, Puja Myles, Richard Branson, Hatim Abdulhussein, Allan Tucker

机构 * Medicine and Healthcare products Regulatory Agency(医疗与健康产品监管局) NHS England(英格兰国家卫生服务)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 47 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16875 2025-05-07 cs.LG 57%

Hybrid Reinforcement Learning and Model Predictive Control for Adaptive Control of Hydrogen-Diesel Dual-Fuel Combustion

Julian Bedei, Murray McBain, Alexander Winkler, Charles Robert Koch, Jakob Andert, David Gordon

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20953 2025-05-07 cs.CL 57%

Clean & Clear: Feasibility of Safe LLM Clinical Guidance

Julia Ive, Felix Jozsa, Nick Jackson, Paulina Bondaronek, Ciaran Scott Hill, Richard Dobson

机构 * University College London(伦敦大学学院) Wolfson Institute of Biomedical Research(生物医学研究沃尔夫森研究所) King’s College Hospital(国王学院医院) National Hospital for Neurology and Neurosurgery(神经病学与神经外科国家医院) King’s College London(伦敦国王学院)

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01542 2025-05-06 cs.HC cs.AI 57%

Emotions in the Loop: A Survey of Affective Computing for Emotional Support

Karishma Hegde, Hemadri Jayalath

机构 * School of Computing University of Georgia(计算学院 佐治亚大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 20 pages, 7 tables, 96 references. Survey paper on affective computing applications using large language models, multimodal AI, and therapeutic chatbots

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16120 2025-04-24 cs.CR cs.AI 57%

A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content

Chaima Njeh, Haïfa Nakouri, Fehmi Jaafar

机构 * Quebec University at Chicoutimi(魁北克大学夏斯库蒂米分校)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments This paper is under revision in the International Journal of Information Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17277 2025-04-24 cs.RO cs.LG 57%

Building Real-time Awareness of Out-of-distribution in Trajectory Prediction for Autonomous Vehicles

Tongfe Guo, Taposh Banerjee, Rui Liu, Lili Su

机构 * Northeastern University(东北大学) University of Pittsburgh(匹兹堡大学) Kent State University(肯特州立大学)

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09137 2025-04-22 cs.CY 57%

Can Large Language Models Become Policy Refinement Partners? Evidence from China's Social Security Studies

Jinghan Ke, Zheng Zhou, Yuxuan Zhao

专题命中 安全训练 :alignment(abstract);分类 cs.CY

Comments 18 pages, 4 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13205 2025-04-21 cs.CR cs.AI 57%

On-Device Watermarking: A Socio-Technical Imperative For Authenticity In The Age of Generative AI

Houssam Kherraz

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI

Comments 10 pages, 3 figures, ICLR 2025, https://openreview.net/forum?id=ygE0U21vxM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11508 2025-04-17 cs.LG 57%

Reward Distance Comparisons Under Transition Sparsity

Clement Nyanhongo, Bruno Miranda Henrique, Eugene Santos

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Published in the TMLR, https://openreview.net/forum?id=haP586YomL

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10873 2025-04-16 cs.CV cs.AI cs.HC 57%

Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles

Tonko E. W. Bossen, Andreas Møgelmose, Ross Greer

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15907 2025-04-16 cs.AI 57%

Belief-State Query Policies for User-Aligned POMDPs

Daniel Bramblett, Siddharth Srivastava

专题命中 安全训练 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08848 2025-04-15 cs.CR cs.AI 57%

X-Guard: Multilingual Guard Agent for Content Moderation

Bibek Upadhayay, Vahid Behzadan, Ph. D

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 34 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01081 2025-04-10 cs.CV cs.CL eess.IV 57%

ShieldGemma 2: Robust and Tractable Image Content Moderation

Wenjun Zeng, Dana Kurniawan, Ryan Mullins, Yuchi Liu, Tamoghna Saha, Dirichi Ike-Njoku, Jindong Gu, Yiwen Song, Cai Xu, Jingjing Zhou, Aparna Joshi, Shravan Dheep, Mani Malek, Hamid Palangi, Joon Baek, Rick Pereira, Karthik Narasimhan

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00441 2025-04-04 cs.CR cs.AI 57%

No Free Lunch with Guardrails

Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02141 2025-04-04 cs.SE cs.AI 57%

On Simulation-Guided LLM-based Code Generation for Safe Autonomous Driving Software

Ali Nouri, Johan Andersson, Kailash De Jesus Hornig, Zhennan Fei, Emil Knabe, Hakan Sivencrona, Beatriz Cabrero-Daniel, Christian Berger

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Accepted in the 29th International Conference on Evaluation and Assessment in Software Engineering (EASE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01719 2025-04-04 cs.LG cs.RO 57%

Beyond Non-Expert Demonstrations: Outcome-Driven Action Constraint for Offline Reinforcement Learning

Ke Jiang, Wen Jiang, Yao Li, Xiaoyang Tan

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18098 2025-04-02 cs.CV cs.LG 57%

Disentangling Safe and Unsafe Corruptions via Anisotropy and Locality

Ramchandran Muthukumar, Ambar Pal, Jeremias Sulam, Rene Vidal

专题命中 安全训练 :alignment(abstract);分类 cs.LG

Comments Published at IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025. Updated Acknowledgements

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13727 2025-04-02 cs.MA cs.AI 57%

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System

Haikuo Du, Fandi Gou, Yunze Cai

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16871 2025-04-01 cs.MA cs.LG stat.ML 57%

Conformal Off-Policy Prediction for Multi-Agent Systems

Tom Kuipers, Renukanandan Tumu, Shuo Yang, Milad Kazemi, Rahul Mangharam, Nicola Paoletti

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted for publication in the 63rd IEEE Conference on Decision and Control (CDC) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21141 2025-03-31 cs.RO cs.LG cs.MA 57%

Safe Human Robot Navigation in Warehouse Scenario

Seth Farrell, Chenghao Li, Hongzhan Yu, Ryo Yoshimitsu, Sicun Gao, Henrik I. Christensen

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21949 2025-03-31 cs.LG 57%

Reward Design for Reinforcement Learning Agents

Rati Devidze

专题命中 安全训练 :alignment(abstract);分类 cs.LG

Comments Doctoral thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20194 2025-03-27 cs.CL 57%

GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization

Zhouhong Gu, Xingzhou Chen, Xiaoran Shi, Tao Wang, Suhang Zheng, Tianyu Li, Hongwei Feng, Yanghua Xiao

专题命中 安全训练 :DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19418 2025-03-26 cs.LG 57%

Multi-Agent Deep Reinforcement Learning for Safe Autonomous Driving with RICS-Assisted MEC

Xueyao Zhang, Bo Yang, Xuelin Cao, Zhiwen Yu, George C. Alexandropoulos, Yan Zhang, Merouane Debbah, Chau Yuen

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17194 2025-03-24 cs.LG 57%

Curriculum RL meets Monte Carlo Planning: Optimization of a Real World Container Management Problem

Abhijeet Pendyala, Tobias Glasmachers

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14563 2025-03-21 cs.SE cs.AI 57%

Workflow for Safe-AI

Suzana Veljanovska, Hans Dermot Doran

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Embedded World Conference, Nuremberg, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00959 2025-03-19 cs.CY 57%

IGGA: A Dataset of Industrial Guidelines and Policy Statements for Generative AIs

Junfeng Jiao, Saleh Afroogh, Kevin Chen, David Atkinson, Amit Dhurandhar

专题命中 安全训练 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03640 2025-03-17 cs.RO cs.LG cs.MA math.OC 57%

Discrete GCBF Proximal Policy Optimization for Multi-agent Safe Optimal Control

Songyuan Zhang, Oswin So, Mitchell Black, Chuchu Fan

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 31 pages, 15 figures; Accepted by the thirteenth International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06550 2025-03-11 cs.CL 57%

BingoGuard: LLM Content Moderation Tools with Risk Levels

Fan Yin, Philippe Laban, Xiangyu Peng, Yilun Zhou, Yixin Mao, Vaibhav Vats, Linnea Ross, Divyansh Agarwal, Caiming Xiong, Chien-Sheng Wu

专题命中 安全训练 :safety(abstract);分类 cs.CL

Comments 10 pages, 4 figures, 4 tables. ICLR 2025 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03911 2025-03-07 cs.RO cs.LG cs.SY eess.SY 57%

Safe LLM-Controlled Robots with Formal Guarantees via Reachability Analysis

Ahmad Hafez, Alireza Naderi Akhormeh, Amr Hegazy, Amr Alanwar

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01334 2025-03-07 cs.RO cs.AI cs.CV cs.HC 57%

A Backbone for Long-Horizon Robot Task Understanding

Xiaoshuai Chen, Wei Chen, Dongmyoung Lee, Yukun Ge, Nicolas Rojas, Petar Kormushev

专题命中 安全训练 :alignment(abstract);分类 cs.AI

Comments 8 pages, 8 figures. This work has been published by IEEE Robotics and Automation Letters (RA-L)

Journal ref IEEE Robotics and Automation Letters, Volume: 10, 2025, 2048 - 2055

详情

展开后加载摘要…

URL PDF HTML 收藏