arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3266 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3266 篇

2508.07075 2025-08-12 cs.LG cs.AI 62%

Surgical Knowledge Rewrite in Compact LLMs: An 'Unlearn-then-Learn' Strategy with ($IA^3$) for Localized Factual Modulation and Catastrophic Forgetting Mitigation

Stanley Ngugi

机构 * Stanley Ngugi(独立研究者)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 2 visual aids

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19250 2025-08-12 cs.LG cs.AI 62%

Robust Behavior Cloning Via Global Lipschitz Regularization

Shili Wu, Yizhao Jin, Puhua Niu, Aniruddha Datta, Sean B. Andersson

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18666 2025-08-01 cs.AI cs.CL 62%

AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents

Haoyu Wang, Christopher M. Poskitt, Jun Sun

机构 * Singapore Management University(新加坡管理大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted by the 48th IEEE/ACM International Conference on Software Engineering (ICSE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23276 2025-07-25 cs.AI cs.CL 62%

Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games

David Guzman Piedrahita, Yongjin Yang, Mrinmaya Sachan, Giorgia Ramponi, Bernhard Schölkopf, Zhijing Jin

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments Published at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14293 2025-07-22 cs.AI cs.CL cs.CV 62%

WebGuard: Building a Generalizable Guardrail for Web Agents

Boyuan Zheng, Zeyi Liao, Scott Salisbury, Zeyuan Liu, Michael Lin, Qinyuan Zheng, Zifan Wang, Xiang Deng, Dawn Song, Huan Sun, Yu Su

机构 * The Ohio State University(俄亥俄州立大学) Scale AI University of California, Berkeley(加州大学伯克利分校)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments We publicly release WebGuard, along with its annotation tools and fine-tuned models, to facilitate open-source research on monitoring and safeguarding web agents. All resources are available at https://github.com/OSU-NLP-Group/WebGuard

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05208 2025-07-17 cs.CR cs.AI cs.LG 62%

Mitigation of Camouflaged Adversarial Attacks in Autonomous Vehicles--A Case Study Using CARLA Simulator

Yago Romano Martinez, Brady Carter, Abhijeet Solanki, Wesam Al Amiri, Syed Rafay Hasan, Terry N. Guo

机构 * Department of Computer Engineering Tennessee Technological University(计算机工程系田纳西技术大学) Department of Computer Science Tennessee Technological University(计算机科学系田纳西技术大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00503 2025-07-09 cs.LG cs.AI cs.RO 62%

Variational OOD State Correction for Offline Reinforcement Learning

Ke Jiang, Wen Jiang, Xiaoyang Tan

机构 * College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics, MIIT Key Laboratory of Pattern Analysis and Machine Intelligence(计算机科学与技术学院,南京航空航天大学,信息科技部模式分析与机器智能重点实验室)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03066 2025-07-08 cs.CL cs.AI 62%

Identification of Potentially Misclassified Crash Narratives using Machine Learning (ML) and Deep Learning (DL)

Sudesh Bhagat, Ibne Farabi Shihab, Jonathan Wood

机构 * Department of Civil, Construction and Environmental Engineering(土木、建设与环境工程系) Iowa State University(爱荷华州立大学) Department of Computer Science(计算机科学系)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23930 2025-07-01 cs.CL cs.AI 62%

Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages

Ruhina Tabasshum Prome, Tarikul Islam Tamiti, Anomadarshi Barua

机构 * Bangladesh Institute of Governance and Management (BIGM)(孟加拉国治理与管理研究所) Department of Cyber Security Engineering, George Mason University(网络安全工程系,乔治·梅森大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06205 2025-06-26 cs.NI cs.AI cs.CL 62%

Leveraging Edge Intelligence and LLMs to Advance 6G-Enabled Internet of Automated Defense Vehicles

Murat Arda Onsu, Poonam Lohan, Burak Kantarci

机构 * School of Electrical Engineering and Computer Science University of Ottawa(电气工程与计算机科学学院渥太华大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments 8 pages, 5 figures, accepted to IEEE Internet of Things Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13060 2025-06-17 cs.AI cs.LG 62%

Rethinking Explainability in the Era of Multimodal AI

Chirag Agarwal

机构 * University of Virginia(弗吉尼亚大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10460 2025-06-17 cs.LO cs.AI cs.LG 62%

Compositional Shielding and Reinforcement Learning for Multi-Agent Systems

Asger Horn Brorholt, Kim Guldstrand Larsen, Christian Schilling

机构 * Aalborg University(奥胡斯大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Journal ref AAMAS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07079 2025-06-10 eess.SY cs.AI cs.LG cs.SY 62%

On the Generalization of Data-Assisted Control in port-Hamiltonian Systems (DAC-pH)

Mostafa Eslami, Maryam Babazadeh

机构 * Sharif University of Technology(谢里夫大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments This paper presents an early investigation of Data-Assisted Control (DAC) with reinforcement learning, showcasing its potential through a simple example. Theoretical analysis is ongoing to establish formal support and guarantees for the proposed approach

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06376 2025-06-10 cs.CL cs.AI 62%

Enhancing Decision-Making of Large Language Models via Actor-Critic

Heng Dong, Kefei Duan, Chongjie Zhang

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments Forty-second International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05113 2025-06-09 cs.CV cs.AI cs.LG 62%

LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Lukas Helff, Felix Friedrich, Manuel Brack, Kristian Kersting, Patrick Schramowski

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments In Proceedings of the 42st International Conference on Machine Learning (ICML 2025), Project page at https://ml-research.github.io/human-centered-genai/projects/llavaguard/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04479 2025-06-06 cs.LG cs.AI 62%

Comparative performance of ensemble models in predicting dental provider types: insights from fee-for-service data

Mohammad Subhi Al-Batah, Muhyeeddin Alqaraleh, Mowafaq Salem Alzboon, Abdullah Alourani

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Journal ref Data and Metadata [Internet]. 2025 Mar. 29 [cited 2025 Jun. 4];4:750

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03614 2025-06-06 cs.CV cs.AI cs.CL cs.CR 62%

VLMs Can Aggregate Scattered Training Patches

Zhanhui Zhou, Lingjie Chen, Chao Yang, Chaochao Lu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23585 2025-06-05 cs.LG cs.CL 62%

On-Policy RL with Optimal Reward Baseline

Yaru Hao, Li Dong, Xun Wu, Shaohan Huang, Zewen Chi, Furu Wei

机构 * Microsoft Research(微软研究院)

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00583 2025-06-03 cs.CL cs.CY cs.HC 62%

The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation

Yuhang Zhou, Yimin Xiao, Wei Ai, Ge Gao

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.CY

Comments 18 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00085 2025-06-03 cs.CL cs.AI 62%

COSMIC: Generalized Refusal Direction Identification in LLM Activations

Vincent Siu, Nicholas Crispino, Zihao Yu, Sam Pan, Zhun Wang, Yang Liu, Dawn Song, Chenguang Wang

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校) University of California, Berkeley(加州大学伯克利分校) University of California, Santa Cruz(加州大学圣克ruz分校)

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments 9 pages, Accepted to ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21841 2025-06-03 cs.LG cs.AI 62%

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints

Jiahui Zhu, Kihyun Yu, Dabeen Lee, Xin Liu, Honghao Wei

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Proceedings of the 41 st International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23933 2025-06-02 cs.LG cs.AI 62%

BIRD: Behavior Induction via Representation-structure Distillation

Galen Pogoncheff, Michael Beyeler

机构 * Department of Computer Science(计算机科学系) UC Santa Barbara(圣巴巴拉大学) Department of Psychological & Brain Sciences(心理学与脑科学系)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22125 2025-05-29 cs.MA cs.AI cs.CY 62%

Sentiment Simulation using Generative AI Agents

Melrose Tia, Jezreel Sophia Lanuzo, Lei Rigi Baltazar, Marie Joy Lopez-Relente, Diwa Malaya Quiñones, Jason Albia

机构 * Netopia AI, Inc.(Netopia AI公司) Institute of Statistics, University of the Philippines Los Baños(菲律宾大学Los Baños统计研究所) Department of Psychology, University of the Philippines Diliman(菲律宾大学Diliman心理学系)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.CY

Comments 18 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15572 2025-05-22 cs.LG cs.AI 62%

Bridging the Domain Gap in Equation Distillation with Reinforcement Feedback

Wangyang Ying, Haoyue Bai, Nanxu Gong, Xinyuan Wang, Sixun Dong, Haifeng Chen, Yanjie Fu

机构 * Arizona State University(亚利桑那州立大学) NEC Laboratories America, Inc.(NEC美国实验室)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06207 2025-05-20 cs.CL cs.AI 62%

Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement

Junyu Lu, Kai Ma, Kaichun Wang, Kelaiti Xiao, Roy Ka-Wei Lee, Bo Xu, Liang Yang, Hongfei Lin

机构 * Key Laboratory of Social Computing and Cognitive Intelligence, Dalian University of Technology(大连理工大学社会计算与认知智能重点实验室) School of Computer Science and Technology, Xinjiang Normal University(新疆师范大学计算机科学与技术学院) Social AI Studio, Singapore University of Technology and Design(新加坡科技设计大学社会AI工作室)

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments 18 pages, accepted at the ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08902 2025-05-15 cs.HC cs.AI cs.CL 62%

Performance Gains of LLMs With Humans in a World of LLMs Versus Humans

Lucas McCullum, Pelagie Ami Agassi, Leo Anthony Celi, Daniel K. Ebner, Chrystinne Oliveira Fernandes, Rachel S. Hicklen, Mkliwa Koumbia, Lisa Soleymani Lehmann, David Restrepo

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16696 2025-05-15 cs.CY cs.AI 62%

Public Constitutional AI

Gilad Abiri

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07911 2025-05-14 cs.LG cs.AI 62%

Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review

Chengmin Zhou, Ville Kyrki, Pasi Fränti, Laura Ruotsalainen

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01619 2025-05-06 cs.LG cs.AI 62%

Skill-based Safe Reinforcement Learning with Risk Planning

Hanping Zhang, Yuhong Guo

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18872 2025-04-29 cs.CL cs.LG 62%

Latent Adversarial Training Improves the Representation of Refusal

Alexandra Abbas, Nora Petrova, Helios Ael Lyons, Natalia Perez-Campanero

机构 * Apart Research(Apart研究)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏