arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3278 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3278 篇

2107.13944 2021-07-30 cs.LG cs.AI stat.ML 62%

Lyapunov-based uncertainty-aware safe reinforcement learning

Ashkan B. Jeddi, Nariman L. Dehghani, Abdollah Shafieezadeh

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Submitted to IEEE Transactions on Neural Networks and Learning Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.11645 2021-07-13 cs.LG cs.AI cs.RO stat.ML 62%

Accelerating Safe Reinforcement Learning with Constraint-mismatched Policies

Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. Ramadge

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments International Conference on Machine Learning (ICML) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.15920 2021-05-19 cs.LG cs.AI cs.RO 62%

Recovery RL: Safe Reinforcement Learning with Learned Recovery Zones

Brijen Thananjeyan, Ashwin Balakrishna, Suraj Nair, Michael Luo, Krishnan Srinivasan, Minho Hwang, Joseph E. Gonzalez, Julian Ibarz, Chelsea Finn, Ken Goldberg

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments RA-L and ICRA 2021. First two authors contributed equally

Journal ref Robotics and Automation Letters (RA-L) and International Conference on Robotics and Automation (ICRA) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.06511 2021-05-17 cs.CL cs.AI 62%

NLP is Not enough -- Contextualization of User Input in Chatbots

Nathan Dolbir, Triyasha Dastidar, Kaushik Roy

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.12558 2021-04-20 cs.AI cs.LG cs.SY eess.SY 62%

Assured Learning-enabled Autonomy: A Metacognitive Reinforcement Learning Framework

Aquib Mustafa, Majid Mazouchi, Subramanya Nageshrao, Hamidreza Modares

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.06390 2021-04-14 cs.CL cs.LG 62%

Detoxifying Language Models Risks Marginalizing Minority Voices

Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, Dan Klein

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

Comments NAACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.01434 2021-04-05 cs.MA cs.AI cs.LG 62%

An Abstraction-based Method to Check Multi-Agent Deep Reinforcement-Learning Behaviors

Pierre El Mqirmi, Francesco Belardinelli, Borja G. León

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Extended version of AAMAS publication under the same name

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.09230 2021-03-17 cs.LG cs.AI cs.RO 62%

Lyapunov Barrier Policy Optimization

Harshit Sikchi, Wenxuan Zhou, David Held

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.03544 2021-03-12 cs.AI cs.CY 62%

Challenges of engineering safe and secure highly automated vehicles

Nadja Marko, Eike Möhlmann, Dejan Ničković, Jürgen Niehaus, Peter Priller, Martijn Rooker

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.04452 2021-03-08 cs.AI cs.LG cs.MA cs.RO 62%

Multi-Agent Safe Planning with Gaussian Processes

Zheqing Zhu, Erdem Bıyık, Dorsa Sadigh

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 5 figures. Published at IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.13045 2021-02-26 cs.LG cs.AI 62%

Iterative Bounding MDPs: Learning Interpretable Policies via Non-Interpretable Methods

Nicholay Topin, Stephanie Milani, Fei Fang, Manuela Veloso

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.12432 2021-02-25 cs.RO cs.AI cs.LG cs.SY eess.SY 62%

Deep Reinforcement Learning for Safe Landing Site Selection with Concurrent Consideration of Divert Maneuvers

Keidai Iiyama, Kento Tomita, Bhavi A. Jagatia, Tatsuwaki Nakagawa, Koki Ho

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 25 pages, 14 figures, This paper is an updated version of Paper AAS 20-583 presented at the AAS/AIAA Astrodynamics Specialist Conference, Online

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.12136 2021-01-22 cs.LG cs.AI cs.RO 62%

Safe Reinforcement Learning via Curriculum Induction

Matteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause, Alekh Agarwal

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.10148 2020-08-25 cs.HC cs.AI cs.LG cs.SY eess.SY 62%

Drive Safe: Cognitive-Behavioral Mining for Intelligent Transportation Cyber-Physical System

Md. Shirajum Munir, Sarder Fakhrul Abedin, Ki Tae Kim, Do Hyeon Kim, Md. Golam Rabiul Alam, Choong Seon Hong

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Submitted to IEEE Transactions on Intelligent Transportation Systems, Special Issue on Technologies for risk mitigation and support of impaired drivers

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.06696 2020-08-19 cs.AI cs.LG cs.RO 62%

Autonomous Braking and Throttle System: A Deep Reinforcement Learning Approach for Naturalistic Driving

Varshit S. Dubey, Ruhshad Kasad, Karan Agrawal

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.06626 2020-08-18 cs.LG cs.AI cs.RO 62%

Safe Reinforcement Learning in Constrained Markov Decision Processes

Akifumi Wachi, Yanan Sui

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 6 figures, Accepted to ICML2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.00089 2020-07-14 cs.LG cs.AI cs.RO stat.ML 62%

Safe, Efficient, and Comfortable Velocity Control based on Reinforcement Learning for Autonomous Driving

Meixin Zhu, Yinhai Wang, Ziyuan Pu, Jingyun Hu, Xuesong Wang, Ruimin Ke

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Under the first-round revision for transportation research part c

Journal ref Transportation Research Part C: Emerging Technologies 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.12069 2020-05-26 cs.LG cs.AI stat.ML 62%

Policy Entropy for Out-of-Distribution Classification

Andreas Sedlmeier, Robert Müller, Steffen Illium, Claudia Linnhoff-Popien

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.02667 2020-05-22 cs.LG cs.AI cs.RO eess.SP 62%

Automated Lane Change Strategy using Proximal Policy Optimization-based Deep Reinforcement Learning

Fei Ye, Xuxin Cheng, Pin Wang, Ching-Yao Chan, Jiucai Zhang

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.08648 2020-04-21 cs.LG cs.AI cs.RO stat.ML 62%

Modeling Survival in model-based Reinforcement Learning

Saeed Moazami, Peggy Doerschuk

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.03237 2020-04-08 cs.LG cs.AI cs.NE 62%

How Do You Act? An Empirical Study to Understand Behavior of Deep Reinforcement Learning Agents

Richard Meyes, Moritz Schneider, Tobias Meisen

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 16 pages, currently under review for publication for the ECMLPKDD 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.00716 2020-04-03 cs.RO cs.AI cs.LG 62%

Constrained-Space Optimization and Reinforcement Learning for Complex Tasks

Ya-Yen Tsai, Bo Xiao, Edward Johns, Guang-Zhong Yang

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in RA-Letters and at ICRA 2020

Journal ref IEEE Robotics and Automation Letters, 5(2) (2020) 682-689

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.12156 2020-03-24 cs.LG cs.AI cs.LO cs.SY eess.SY stat.ML 62%

Cautious Reinforcement Learning with Logical Constraints

Mohammadhosein Hasanbeig, Alessandro Abate, Daniel Kroening

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to AAMAS 2020. arXiv admin note: text overlap with arXiv:1902.00778

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.11699 2020-02-12 cs.RO cs.AI cs.LG cs.MA 62%

Multi-Vehicle Mixed-Reality Reinforcement Learning for Autonomous Multi-Lane Driving

Rupert Mitchell, Jenny Fletcher, Jacopo Panerati, Amanda Prorok

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.02390 2019-11-25 cs.LG cs.CY stat.ML 62%

Migration through Machine Learning Lens -- Predicting Sexual and Reproductive Health Vulnerability of Young Migrants

Amber Nigam, Pragati Jaiswal, Uma Girkar, Teertha Arora, Leo A. Celi

专题命中 安全训练 :safety(abstract);分类 cs.CY、cs.LG

Comments Accepted for Machine Learning for Health (ML4H) at NeurIPS 2019 - Extended Abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.13726 2019-10-31 cs.LG cs.AI cs.RO stat.ML 62%

Safe Exploration for Interactive Machine Learning

Matteo Turchetta, Felix Berkenkamp, Andreas Krause

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.12189 2019-07-02 eess.SY cs.AI cs.LG cs.SY 62%

Learning-based Model Predictive Control for Safe Exploration and Reinforcement Learning

Torsten Koller, Felix Berkenkamp, Matteo Turchetta, Joschka Boedecker, Andreas Krause

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 14 pages, 7 figures. arXiv admin note: substantial text overlap with arXiv:1803.08287

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.07705 2019-07-02 cs.RO cs.AI cs.LG 62%

Multi-Objective Autonomous Braking System using Naturalistic Dataset

Rafael Vasquez, Bilal Farooq

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted in the proceedings of IEEE ITSC2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.04127 2019-05-13 cs.LG cs.AI 62%

Design of Artificial Intelligence Agents for Games using Deep Reinforcement Learning

Andrei Claudiu Roibu

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments Dissertation submitted to the University of Sheffield in partial fulfilment of the requirements for the degree of Master of Engineering. 98 pages, 21 Tables, 58 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.01303 2019-05-07 cs.LG cs.AI cs.MA stat.ML 62%

Autonomous Air Traffic Controller: A Deep Multi-Agent Reinforcement Learning Approach

Marc Brittain, Peng Wei

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏