arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3278 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3278 篇

2505.08902 2025-05-15 cs.HC cs.AI cs.CL 62%

Performance Gains of LLMs With Humans in a World of LLMs Versus Humans

Lucas McCullum, Pelagie Ami Agassi, Leo Anthony Celi, Daniel K. Ebner, Chrystinne Oliveira Fernandes, Rachel S. Hicklen, Mkliwa Koumbia, Lisa Soleymani Lehmann, David Restrepo

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16696 2025-05-15 cs.CY cs.AI 62%

Public Constitutional AI

Gilad Abiri

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07911 2025-05-14 cs.LG cs.AI 62%

Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review

Chengmin Zhou, Ville Kyrki, Pasi Fränti, Laura Ruotsalainen

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01619 2025-05-06 cs.LG cs.AI 62%

Skill-based Safe Reinforcement Learning with Risk Planning

Hanping Zhang, Yuhong Guo

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18872 2025-04-29 cs.CL cs.LG 62%

Latent Adversarial Training Improves the Representation of Refusal

Alexandra Abbas, Nora Petrova, Helios Ael Lyons, Natalia Perez-Campanero

机构 * Apart Research(Apart研究)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10318 2025-04-29 cs.LG cs.AI 62%

Enhance Exploration in Safe Reinforcement Learning with Contrastive Representation Learning

Duc Kien Doan, Bang Giang Le, Viet Cuong Ta

机构 * HMI Laboratory VNU University of Engineering and Technology(VNU工程技术大学HMI实验室)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at ACIIDS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13969 2025-04-24 cs.HC cs.AI cs.CY 62%

Tinker Tales: Interactive Storytelling Framework for Early Childhood Narrative Development and AI Literacy

Nayoung Choi, Peace Cyebukayire, Jinho D. Choi

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15425 2025-04-23 cs.RO cs.AI cs.LG cs.MA math.OC 62%

Solving Multi-Agent Safe Optimal Control with Distributed Epigraph Form MARL

Songyuan Zhang, Oswin So, Mitchell Black, Zachary Serlin, Chuchu Fan

机构 * MIT(麻省理工学院) MIT Lincoln Laboratory(麻省理工学院林肯实验室)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 28 pages, 16 figures; Accepted by Robotics: Science and Systems 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15369 2025-04-23 cs.LG cs.AI cs.RO 62%

Solving New Tasks by Adapting Internet Video Knowledge

Calvin Luo, Zilai Zeng, Yilun Du, Chen Sun

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments ICLR 2025. Project Webpage: https://diffusion-supervision.github.io/adapt2act/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17070 2025-04-18 cs.CR cs.CL cs.LG 62%

Contextual Agent Security: A Policy for Every Purpose

Lillian Tsai, Eugene Bagdasarian

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

Comments Workshop in Hot Topics in Operating Systems (HotOS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12299 2025-04-17 cs.AI cs.CV cs.LG 62%

Adapting a World Model for Trajectory Following in a 3D Game

Marko Tot, Shu Ishida, Abdelhak Lemkhenter, David Bignell, Pallavi Choudhury, Chris Lovett, Luis França, Matheus Ribeiro Furtado de Mendonça, Tarun Gupta, Darren Gehring, Sam Devlin, Sergio Valcarcel Macua, Raluca Georgescu

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12082 2025-04-17 cs.CL cs.AI 62%

Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection

Yumin Kim, Hwanhee Lee

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12808 2025-04-11 cs.AI cs.CY cs.HC 62%

Conversational Medical AI: Ready for Practice

Antoine Lizée, Pierre-Auguste Beaucoté, James Whitbeck, Marion Doumeingts, Anaël Beaugnon, Isabelle Feldhaus

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY

Comments Accepted to AAAI25 (Oral, workshop) 14 pages, 7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04215 2025-04-08 cs.CL cs.AI 62%

Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability

Vishnu Kabir Chhabra, Mohammad Mahdi Khalili

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21504 2025-03-28 cs.CL cs.AI cs.CV 62%

Keyword-Oriented Multimodal Modeling for Euphemism Identification

Yuxue Hu, Junsong Li, Meixuan Chen, Dongyu Su, Tongguan Wang, Ying Sha

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07671 2025-03-26 stat.ML cs.AI cs.LG 62%

Probabilistic Shielding for Safe Reinforcement Learning

Edwin Hamel-De le Court, Francesco Belardinelli, Alexander W. Goodall

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 3 figures, Conference: AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12198 2025-03-25 cs.CL cs.CV cs.LG 62%

Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations

Naquee Rizwan, Paramananda Bhaskar, Mithun Das, Swadhin Satyaprakash Majhi, Punyajoy Saha, Animesh Mukherjee

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20089 2025-03-21 cs.LG cs.CL cs.CR 62%

Robust LLM safeguarding via refusal feature adversarial training

Lei Yu, Virginie Do, Karen Hambardzumyan, Nicola Cancedda

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14703 2025-03-20 cs.CL cs.AI 62%

Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics

Seungbeen Lee, Seungwon Lim, Seungju Han, Giyeong Oh, Hyungjoo Chae, Jiwan Chung, Minju Kim, Beong-woo Kwak, Yeonsoo Lee, Dongha Lee, Jinyoung Yeo, Youngjae Yu

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to NAACL2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12753 2025-03-18 cs.NI cs.AI cs.LG 62%

SafeSlice: Enabling SLA-Compliant O-RAN Slicing via Safe Deep Reinforcement Learning

Ahmad M. Nagib, Hatem Abou-Zeid, Hossam S. Hassanein

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments This article has been accepted for presentation in the IEEE International Conference on Machine Learning for Communication and Networking (ICMLCN) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11696 2025-03-18 eess.SY cs.AI cs.LG cs.SY 62%

Balancing SoC in Battery Cells using Safe Action Perturbations

E Harshith Kumar Yadav, Rahul Narava, Anshika, Shashi Shekher Jha

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13187 2025-03-11 cs.LG cs.AI cs.RO 62%

A Survey of Sim-to-Real Methods in RL: Progress, Prospects and Challenges with Foundation Models

Longchao Da, Justin Turnau, Thirulogasankar Pranav Kutralingam, Alvaro Velasquez, Paulo Shakarian, Hua Wei

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 19 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10998 2025-03-06 eess.SY cs.AI cs.LG cs.LO cs.SY 62%

Provably Safe Neural Network Controllers via Differential Dynamic Logic

Samuel Teuber, Stefan Mitsch, André Platzer

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 39 pages (main paper has 10 pages), 13 figures; Accepted at the Thirty-Eighth Annual Conference on Neural Information Processing Systems (NeurIPS 2024)

Journal ref in Advances in Neural Information Processing Systems, 2024, pp. 1586-1624

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19271 2025-02-28 cs.CL cs.AI cs.IR 62%

AutoPureData: Automated Filtering of Undesirable Web Data to Update LLM Knowledge

Praneeth Vadlapati

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments Final version

Journal ref Journal of Mathematical & Computer Applications, 3 (2024) E121

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12411 2025-02-19 cs.CL cs.AI 62%

Gradient Co-occurrence Analysis for Detecting Unsafe Prompts in Large Language Models

Jingyuan Yang, Bowen Yan, Rongjun Li, Ziyu Zhou, Xin Chen, Zhiyong Feng, Wei Peng

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14853 2025-02-18 cs.CL cs.AI cs.CR 62%

Atoxia: Red-teaming Large Language Models with Target Toxic Answers

Yuhao Du, Zhuo Li, Pengyu Cheng, Xiang Wan, Anningzhe Gao

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to Findings of NAACL-2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10431 2025-02-18 cs.LG cs.AI 62%

Leveraging Constraint Violation Signals For Action-Constrained Reinforcement Learning

Janaka Chathuranga Brahmanage, Jiajing Ling, Akshat Kumar

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 7 pages and 5 pages supplementary

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09886 2025-02-17 cs.RO cs.AI cs.LG 62%

Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos

Weirui Ye, Fangchen Liu, Zheng Ding, Yang Gao, Oleh Rybkin, Pieter Abbeel

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19471 2025-02-17 cs.RO cs.AI cs.CL cs.FL 62%

SELP: Generating Safe and Efficient Task Plans for Robot Agents with Large Language Models

Yi Wu, Zikang Xiong, Yiran Hu, Shreyash S. Iyengar, Nan Jiang, Aniket Bera, Lin Tan, Suresh Jagannathan

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted for presentation at the 2025 IEEE International Conference on Robotics and Automation (ICRA), May 19-23, 2025, Atlanta, USA, and for inclusion in the conference proceeding

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04184 2025-02-17 cs.LO cs.AI cs.LG cs.RO 62%

Shield Synthesis for LTL Modulo Theories

Andoni Rodriguez, Guy Amir, Davide Corsi, Cesar Sanchez, Guy Katz

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments To appear in AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏