arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3281 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3281 篇

2302.09491 2023-02-21 cs.CR cs.AI cs.CV 57%

X-Adv: Physical Adversarial Object Attacks against X-ray Prohibited Item Detection

Aishan Liu, Jun Guo, Jiakai Wang, Siyuan Liang, Renshuai Tao, Wenbo Zhou, Cong Liu, Xianglong Liu, Dacheng Tao

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Accepted by USENIX Security 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04321 2023-02-16 cs.RO cs.AI 57%

Shared Information-Based Safe And Efficient Behavior Planning For Connected Autonomous Vehicles

Songyang Han, Shanglin Zhou, Lynn Pepin, Jiangwei Wang, Caiwen Ding, Fei Miao

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments This paper gets the Best Paper Award in the DCAA workshop of AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02788 2023-02-07 cs.LG 57%

A Strong Baseline for Batch Imitation Learning

Matthew Smith, Lucas Maystre, Zhenwen Dai, Kamil Ciosek

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 28 pages (10 main, 18 appendix), 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05838 2023-01-31 cs.CV cs.AI cs.HC 57%

(Safe) SMART Hands: Hand Activity Analysis and Distraction Alerts Using a Multi-Camera Framework

Ross Greer, Lulua Rakla, Anish Gopalan, Mohan Trivedi

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00904 2023-01-04 cs.RO cs.LG cs.SY eess.SY 57%

Safe Reinforcement Learning for an Energy-Efficient Driver Assistance System

Habtamu Hailemichael, Beshah Ayalew, Lindsey Kerbel, Andrej Ivanco, Keith Loiselle

专题命中 安全训练 :safety(abstract);分类 cs.LG

Journal ref IFAC-PapersOnLine, Volume 55, Issue 37, 2022, Pages 615-620

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.13819 2022-12-29 cs.AI 57%

Don't do it: Safer Reinforcement Learning With Rule-based Guidance

Ekaterina Nikonova, Cheng Xue, Jochen Renz

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.13769 2022-12-29 cs.LG 57%

Lexicographic Multi-Objective Reinforcement Learning

Joar Skalse, Lewis Hammond, Charlie Griffin, Alessandro Abate

专题命中 安全训练 :safety(abstract);分类 cs.LG

Journal ref IJCAI 2022; Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence. Main Track, Pages 3430-3436

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.11746 2022-12-23 cs.LG cs.MA 57%

Certified Policy Smoothing for Cooperative Multi-Agent Reinforcement Learning

Ronghui Mu, Wenjie Ruan, Leandro Soriano Marcolino, Gaojie Jin, Qiang Ni

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments This paper will appear in AAAI2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.10904 2022-12-16 cs.LG stat.ML 57%

Reward Shaping for Human Learning via Inverse Reinforcement Learning

Mark A. Rucker, Layne T. Watson, Matthew S. Gerber, Laura E. Barnes

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments This paper has been modified considerably for resubmission to Journal of Machine Learning Research, for source code, see https://github.com/mrucker/kpirl-kla

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02010 2022-12-06 cs.MA cs.AI cs.GT 57%

Multi Agent Path Finding using Evolutionary Game Theory

Sheryl Paul, Jyotirmoy V. Deshmukh

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00639 2022-12-05 cs.LG cs.DC 57%

Launchpad: Learning to Schedule Using Offline and Online RL Methods

Vanamala Venkataswamy, Jake Grigsby, Andrew Grimshaw, Yanjun Qi

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00669 2022-12-02 physics.geo-ph cs.AI 57%

A POMDP Model for Safe Geological Carbon Sequestration

Anthony Corso, Yizheng Wang, Markus Zechner, Jef Caers, Mykel J. Kochenderfer

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Accepted at NeurIPS 2022 Workshop on Tackling Climate Change with Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11056 2022-11-22 cs.RO cs.AI cs.SY eess.SY 57%

Safe Control Under Input Limits with Neural Control Barrier Functions

Simin Liu, Changliu Liu, John Dolan

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments CORL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11027 2022-11-22 cs.RO cs.LG cs.SY eess.SY 57%

Safe Reinforcement Learning using Data-Driven Predictive Control

Mahmoud Selim, Amr Alanwar, M. Watheq El-Kharashi, Hazem M. Abbas, Karl H. Johansson

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03032 2022-11-08 cs.LG 57%

Decentralized Policy Optimization

Kefan Su, Zongqing Lu

专题命中 安全训练 :DPO(abstract);分类 cs.LG

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.02131 2022-11-07 cs.RO cs.LG 57%

Safe Real-World Autonomous Driving by Learning to Predict and Plan with a Mixture of Experts

Stefano Pini, Christian S. Perone, Aayush Ahuja, Ana Sofia Rufino Ferreira, Moritz Niendorf, Sergey Zagoruyko

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.13477 2022-10-13 cs.AI 57%

Parametrically Retargetable Decision-Makers Tend To Seek Power

Alexander Matt Turner, Prasad Tadepalli

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 10-page main paper, 36 pages total, poster at NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07870 2022-10-12 cs.AI 57%

How to talk so AI will learn: Instructions, descriptions, and autonomy

Theodore R Sumers, Robert D Hawkins, Mark K Ho, Thomas L Griffiths, Dylan Hadfield-Menell

专题命中 安全训练 :alignment(abstract);分类 cs.AI

Comments 10 pages, 5 figures. Published as a conference paper at NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02552 2022-10-07 cs.LG 57%

Towards Safe Mechanical Ventilation Treatment Using Deep Offline Reinforcement Learning

Flemming Kondrup, Thomas Jiralerspong, Elaine Lau, Nathan de Lara, Jacob Shkrob, My Duc Tran, Doina Precup, Sumana Basu

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments to be published in IAAI (Innovative Applications of Artificial Intelligence) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.04749 2022-10-04 cs.RO cs.LG 57%

Bilateral Deep Reinforcement Learning Approach for Better-than-human Car Following Model

Tianyu Shi, Yifei Ai, Omar ElSamadisy, Baher Abdulhai

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00755 2022-08-24 cs.AI 57%

Safe Reinforcement Learning via Shielding under Partial Observability

Steven Carr, Nils Jansen, Sebastian Junges, Ufuk Topcu

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 21 pages, 28 Figures, 3 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.08409 2022-08-03 cs.LG 57%

How to Learn from Risk: Explicit Risk-Utility Reinforcement Learning for Efficient and Safe Driving Strategies

Lukas M. Schmidt, Sebastian Rietsch, Axel Plinge, Bjoern M. Eskofier, Christopher Mutschler

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.13446 2022-07-28 cs.LG cs.FL 57%

Dynamic Shielding for Reinforcement Learning in Black-Box Environments

Masaki Waga, Ezequiel Castellano, Sasinee Pruekprasert, Stefan Klikovits, Toru Takisaka, Ichiro Hasuo

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments This is the author (and extended) version of the manuscript of the same name published in the proceedings of the 20th International Symposium on Automated Technology for Verification and Analysis (ATVA 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.12674 2022-07-19 cs.LG cs.RO 57%

MetaDrive: Composing Diverse Driving Scenarios for Generalizable Reinforcement Learning

Quanyi Li, Zhenghao Peng, Lan Feng, Qihang Zhang, Zhenghai Xue, Bolei Zhou

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Source code, documentation, and demo video are available at https://metadriverse.github.io/metadrive . More research projects based on MetaDrive simulator are listed at https://metadriverse.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10158 2022-07-05 cs.LG cs.MA 57%

Certifiably Robust Policy Learning against Adversarial Communication in Multi-agent Systems

Yanchao Sun, Ruijie Zheng, Parisa Hassanzadeh, Yongyuan Liang, Soheil Feizi, Sumitra Ganesh, Furong Huang

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10797 2022-06-23 cs.LG cs.CV cs.RO 57%

Imitation Learning for Generalizable Self-driving Policy with Sim-to-real Transfer

Zoltán Lőrincz, Márton Szemenyei, Róbert Moni

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted by ICLR 2022 Workshop on Generalizable Policy Learning in Physical World. Source code is available at: https://github.com/lzoltan35/duckietown_imitation_learning

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.03231 2022-06-20 cs.LG stat.ML 57%

Smoothing Policies and Safe Policy Gradients

Matteo Papini, Matteo Pirotta, Marcello Restelli

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.10658 2022-06-09 cs.MA cs.LG cs.RO cs.SY eess.SY 57%

Decentralized Safe Multi-agent Stochastic Optimal Control using Deep FBSDEs and ADMM

Marcus A. Pereira, Augustinos D. Saravanos, Oswin So, Evangelos A. Theodorou

专题命中 安全训练 :safety(abstract);分类 cs.LG

Journal ref Robotics: Science and Systems (RSS), 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.12494 2022-05-20 cs.LG cs.RO stat.ML 57%

SEMI: Self-supervised Exploration via Multisensory Incongruity

Jianren Wang, Ziwen Zhuang, Hang Zhao

专题命中 安全训练 :alignment(abstract);分类 cs.LG

Comments Accepted at ICRA 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08657 2022-05-19 cs.RO cs.AI cs.HC 57%

Intuitive and Efficient Human-robot Collaboration via Real-time Approximate Bayesian Inference

Javier Felip Leon, David Gonzalez-Aguirre, Lama Nachman

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏