arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3266 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3266 篇

2412.15462 2024-12-23 cs.RO cs.AI cs.CL cs.HC cs.LG 67%

TalkWithMachines: Enhancing Human-Robot Interaction for Interpretable Industrial Robotics Through Large/Vision Language Models

Ammar N. Abbas, Csaba Beleznai

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments This paper has been accepted for publication in the proceedings of the 2024 Eighth IEEE International Conference on Robotic Computing (IRC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07933 2024-11-01 cs.CL cs.AI cs.LG 67%

Large Language Model Unlearning via Embedding-Corrupted Prompts

Chris Yuhao Liu, Yaxuan Wang, Jeffrey Flanigan, Yang Liu

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2024 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15762 2024-10-24 cs.LG cs.AI cs.CL 67%

Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning

Kaiwen Wang, Rahul Kidambi, Ryan Sullivan, Alekh Agarwal, Christoph Dann, Andrea Michi, Marco Gelmi, Yunxuan Li, Raghav Gupta, Avinava Dubey, Alexandre Ramé, Johan Ferret, Geoffrey Cideron, Le Hou, Hongkun Yu, Amr Ahmed, Aranyak Mehta, Léonard Hussenot, Olivier Bachem, Edouard Leurent

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 40 pages. Findings of EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01787 2024-08-19 cs.CY cs.AI cs.LG 67%

Harm Amplification in Text-to-Image Models

Susan Hao, Renee Shelby, Yuchi Liu, Hansa Srinivasan, Mukul Bhutani, Burcu Karagol Ayan, Ryan Poplin, Shivani Poddar, Sarah Laszlo

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01452 2024-08-06 cs.CY cs.AI cs.LG 67%

Building a Domain-specific Guardrail Model in Production

Mohammad Niknazar, Paul V Haley, Latha Ramanan, Sang T. Truong, Yedendra Shrinivasan, Ayan Kumar Bhowmick, Prasenjit Dey, Ashish Jagmohan, Hema Maheshwari, Shom Ponoth, Robert Smith, Aditya Vempaty, Nick Haber, Sanmi Koyejo, Sharad Sundararajan

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16900 2024-07-25 cs.LG cs.AI cs.CY 67%

Regulating AI Adaptation: An Analysis of AI Medical Device Updates

Kevin Wu, Eric Wu, Kit Rodolfa, Daniel E. Ho, James Zou

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Journal ref CHIL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16268 2024-07-19 cs.LG cs.AI cs.CY 67%

Foundation Model Transparency Reports

Rishi Bommasani, Kevin Klyman, Shayne Longpre, Betty Xiong, Sayash Kapoor, Nestor Maslej, Arvind Narayanan, Percy Liang

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Journal ref Published in AIES 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17806 2024-06-27 cs.CL cs.AI cs.CR cs.CV cs.LG 67%

MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?

Xirui Li, Hengguang Zhou, Ruochen Wang, Tianyi Zhou, Minhao Cheng, Cho-Jui Hsieh

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11253 2024-06-26 cs.LG cs.AI cs.CL 67%

Aligning Large Language Models by On-Policy Self-Judgment

Sangkyu Lee, Sungdong Kim, Ashkan Yousefpour, Minjoon Seo, Kang Min Yoo, Youngjae Yu

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published as a main conference paper at ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15447 2024-05-07 cs.CL cs.AI cs.CR cs.LG 67%

Are aligned neural networks adversarially aligned?

Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, Ludwig Schmidt

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03683 2024-04-08 cs.LG cs.AI cs.CL 67%

Stream of Search (SoS): Learning to Search in Language

Kanishk Gandhi, Denise Lee, Gabriel Grand, Muxin Liu, Winson Cheng, Archit Sharma, Noah D. Goodman

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08461 2024-04-02 cs.CL cs.AI cs.LG 67%

DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Yongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon, Afshin Rostamizadeh, Sanjiv Kumar, Jean-François Kagy, Rishabh Agarwal

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13649 2024-01-18 cs.LG cs.AI cs.CL 67%

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem

专题命中 安全训练 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ICLR 2024. First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03186 2023-12-07 cs.LG cs.AI cs.CY 67%

Data-Driven Traffic Reconstruction and Kernel Methods for Identifying Stop-and-Go Congestion

Edgar Ramirez Sanchez, Shreyaa Raghavan, Cathy Wu

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Presented at NeurIPS 2023 workshops: Tackling Climate Change with Machine Learning & Computational Sustainability

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14540 2023-09-27 cs.LG cs.AI cs.CV cs.CY 67%

Effect of roundabout design on the behavior of road users: A case study of roundabouts with application of Unsupervised Machine Learning

Tasnim M. Dwekat, Ayda A. Almsre, Huthaifa I. Ashqar

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13816 2023-07-21 cs.LG cs.AI cs.CL cs.PL 67%

Execution-based Code Generation using Deep Reinforcement Learning

Parshin Shojaee, Aneesh Jain, Sindhu Tipirneni, Chandan K. Reddy

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published in Transactions on Machine Learning Research (TMLR), 2023

Journal ref Transactions on Machine Learning Research (TMLR), 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06609 2023-03-07 cs.RO 67%

TrafficGen: Learning to Generate Diverse and Realistic Traffic Scenarios

Lan Feng, Quanyi Li, Zhenghao Peng, Shuhan Tan, Bolei Zhou

专题命中 安全训练 :safety(abstract);AI safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.07495 2020-06-16 cs.NE 67%

Open Questions in Creating Safe Open-ended AI: Tensions Between Control and Creativity

Adrien Ecoffet, Jeff Clune, Joel Lehman

专题命中 安全训练 :safety(abstract);AI safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20792 2025-09-26 cs.CV cs.AI cs.LG 66%

DAC-LoRA: Dynamic Adversarial Curriculum for Efficient and Robust Few-Shot Adaptation

Ved Umrajkar

机构 * Indian Institute of Technology, Roorkee(印度理工学院拉胡尔分校)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG;trustworthy(comments)

Comments Accepted at ICCV2025 Workshop on Safe and Trustworthy Multimodal AI Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04867 2025-08-28 cs.AI cs.LG 66%

Think Smart, Act SMARL! Analyzing Probabilistic Logic Shields for Multi-Agent Reinforcement Learning

Satchit Chatterji, Erman Acar

机构 * IvI \& ILLC, University of Amsterdam

专题命中 安全训练 :safety(abstract,comments);分类 cs.AI、cs.LG

Comments Accepted to the 28th European Conference on Artificial Intelligence (ECAI 2025) --- 21 pages, 15 figures, Earlier title: "Analyzing Probabilistic Logic Driven Safety in Multi-Agent Reinforcement Learning"; (changed for specificity and clarity)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14110 2024-10-30 cs.LG cs.AI 66%

When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective

Hao Sun, Alex J. Chan, Nabeel Seedat, Alihan Hüyük, Mihaela van der Schaar

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG;RLHF(comments)

Comments Reward Modeling, Large Language Models, RLHF, Off-Policy Evaluation, Data-Centric AI, Data-Centric Reinforcement Learning, Reinforcement Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16209 2024-10-29 cs.LG cs.AI 66%

C-MCTS: Safe Planning with Monte Carlo Tree Search

Dinesh Parthasarathy, Georgios Kontes, Axel Plinge, Christopher Mutschler

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG;trustworthy(comments)

Comments Workshop on Safe & Trustworthy Agents @NeurIPS2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09976 2024-07-02 cs.LG cs.AI 66%

Robust Model-Based Reinforcement Learning with an Adversarial Auxiliary Model

Siemen Herremans, Ali Anwar, Siegfried Mercelis

专题命中 安全训练 :safety(abstract,comments);分类 cs.AI、cs.LG

Comments Will be presented at the RL Safety Workshop at RLC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.04472 2019-12-11 cs.LG cs.AI stat.ML 66%

Deep Bayesian Reward Learning from Preferences

Daniel S. Brown, Scott Niekum

专题命中 安全训练 :safety(abstract,comments);分类 cs.AI、cs.LG

Comments Workshop on Safety and Robustness in Decision Making at the 33rd Conference on Neural Information Processing Systems (NeurIPS) 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16386 2026-08-18 cs.CL cs.LG 新提交 62%

Mint-Agent: Introducing Finance-Native Agentic Foundation Models

Mint-Agent:引入金融原生智能体基础模型

Agent Team, B. Zhang, Yaze Geng, Lei Tang, Yaoyang Yi, Zonghan Wu, Yifan Hu, Kun Wang, Qingsong Wen, Yilei Shao

专题命中 安全训练 :trustworthy(abstract);分类 cs.CL、cs.LG

AI总结 本研究提出金融原生智能体模型Mint-Agent,通过三大支柱开发出Mint-Cu(9B)和Mint-Ag(27B),在多个金融基准测试中展现出优异的可靠性与可执行性,为可信金融智能提供了新路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15673 2026-08-18 cs.LG cs.AI 新提交 62%

PL-Guard: Probabilistic Logic Reasoning for LLM Guardrails

PL-Guard:面向大语言模型安全护栏的概率逻辑推理

Satchit Chatterji, Shihan Wang, Giovanni Sileno, Erman Acar

机构 * University of Amsterdam(阿姆斯特丹大学) Utrecht University(乌得勒支大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

AI总结 该研究提出神经符号安全护栏架构PL-Guard,通过分离神经落地与概率符号推理,在XSTest基准上大幅降低大语言模型的不安全依从率,虽过度拒绝率略高,但提升了推理可审计性。

Comments Preliminary version of this paper was presented at the IJCAI 2026 Workshop on Logical and Symbolic Reasoning of Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13160 2026-08-14 cs.CL cs.AI 新提交 62%

Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering

更好的分解,自由聚合:面向多语言多跳问答的合成器折叠框架

Yilin Wang, Yuchun Fan, Weidong Bao, Zili Wei, Shi Feng, Tong Xiao, Zhengtao Yu, Jingbo Zhu

机构 * School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院) Kunming University of Science and Technology(昆明理工大学)

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 针对多语言多跳问答中现有方法的翻译噪声与错误放大问题,提出Syfer框架,通过推迟翻译、格式受限分解与质量检查实现性能与成本的良好平衡。

Comments Accepted by NLPCC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06564 2026-08-14 cs.LG cs.CL 版本更新 62%

Which Decisions Low-Bit Quantization Breaks, and How to Predict Them

量化损伤是乘法性的,而非加法性的

Zekun Wu, Swati Dhiman, Adriano Koshiyama

机构 * Holistic AI(霍利斯提克人工智能公司) University College London(伦敦大学学院)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

AI总结 该研究发现大型语言模型的量化损伤是乘法性的,而非加法性的,提出边际收缩概念,拟合关系可准确预测决策翻转率,增加1位是修复损伤最廉价的方式。

Comments 18 pages, 9 figures, 8 tables. Under review at the Third Workshop on Uncertainty-Aware NLP (UncertaiNLP), EMNLP 2026 (non-archival)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11658 2026-08-13 cs.LG cs.AI cs.MA 新提交 62%

Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning

智能体策略组合是否安全?重新思考合作多智能体强化学习中的后继特征迁移

Zijian Zhao, Sen Li

机构 * The Hong Kong University of Science and Technology(香港科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

AI总结 针对多智能体强化学习中独立策略组合不安全的问题,提出MA-USFA分层方法,兼顾策略组合的安全性与灵活性,无需任务适配即可部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11505 2026-08-11 cs.LG cs.AI 版本更新 62%

Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update

代理探索与可重用引导:基于代理引导更新信号的模块化大语言模型训练后范式

Daocheng Fu, Rong Wu, Yu Yang, Jianbiao Mei, Licheng Wen, Pinlong Cai, Xuemeng Yang, Yong Liu, Botian Shi, Yu Qiao

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 研究针对大语言模型训练后优化信号耦合问题,提出PUST框架,通过代理模型探索、提取并传输相对改进信号,解耦更新信号探索与分布对齐,减少计算开销,支持异步处理与跨模型转移,提升训练后范式的模块化、可重用性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏