arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3281 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3281 篇

2201.10436 2022-05-13 cs.AI 57%

Safe AI -- How is this Possible?

Harald Rueß, Simon Burton

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 42 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.14255 2022-05-03 cs.AI cs.MA 57%

Human-in-the-loop online multi-agent approach to increase trustworthiness in ML models through trust scores and data augmentation

Gusseppe Bravo-Rocca, Peini Liu, Jordi Guitart, Ajay Dholakia, David Ellison, Miroslav Hodak

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Extended version of short paper accepted at IEEE COMPSAC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00176 2022-05-03 cs.CL 57%

Building a Role Specified Open-Domain Dialogue System Leveraging Large-Scale Language Models

Sanghwan Bae, Donghyun Kwak, Sungdong Kim, Donghoon Ham, Soyoung Kang, Sang-Woo Lee, Woomyoung Park

专题命中 安全训练 :safety(abstract);分类 cs.CL

Comments Accepted to NAACL2022 as a long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.01977 2022-04-26 cs.AI 57%

Safe RAN control: A Symbolic Reinforcement Learning Approach

Alexandros Nikou, Anusha Mujumdar, Vaishnavi Sundararajan, Marin Orlic, Aneta Vulgarakis Feljan

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments To appear in International Conference of Control and Automation (ICCA) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10451 2022-04-25 eess.SY cs.AR cs.LG cs.PF cs.SY 57%

SCOPE: Safe Exploration for Dynamic Computer Systems Optimization

Hyunji Kim, Ahsan Pervaiz, Henry Hoffmann, Michael Carbin, Yi Ding

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.06835 2022-04-15 cs.RO cs.AI 57%

GloCAL: Glocalized Curriculum-Aided Learning of Multiple Tasks with Application to Robotic Grasping

Anil Kurkcu, Cihan Acar, Domenico Campolo, Keng Peng Tee

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.10698 2022-03-22 cs.CR cs.AI 57%

A Policy Driven AI-Assisted PoW Framework

Trisha Chakraborty, Shaswata Mitra, Sudip Mittal, Maxwell Young

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.05938 2022-03-07 cs.CV cs.LG cs.SE 57%

RGB cameras failures and their effects in autonomous driving applications

Francesco Secci, Andrea Ceccarelli

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.00369 2022-03-02 cs.RO cs.LG 57%

Approximating a deep reinforcement learning docking agent using linear model trees

Vilde B. Gjærum, Ella-Lovise H. Rørvik, Anastasios M. Lekkas

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.07789 2022-02-17 cs.LG 57%

Safe Reinforcement Learning by Imagining the Near Future

Garrett Thomas, Yuping Luo, Tengyu Ma

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted at NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02793 2022-02-11 cs.AI cs.MA 57%

Multi-Agent Constrained Policy Optimisation

Shangding Gu, Jakub Grudzien Kuba, Munning Wen, Ruiqing Chen, Ziyan Wang, Zheng Tian, Jun Wang, Alois Knoll, Yaodong Yang

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01388 2022-02-02 cs.LG cs.RO cs.SY eess.SY 57%

Neural Network Verification in Control

Michael Everett

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments arXiv admin note: text overlap with arXiv:2108.04140, arXiv:2004.06496

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.09101 2022-01-14 eess.SY cs.LG cs.SY math.OC 57%

Safe Policies for Reinforcement Learning via Primal-Dual Methods

Santiago Paternain, Miguel Calvo-Fullana, Luiz F. O. Chamon, Alejandro Ribeiro

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments arXiv admin note: text overlap with arXiv:1910.13393

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.13827 2022-01-11 cs.LG 57%

Learning to Simulate Self-Driven Particles System with Coordinated Policy Optimization

Zhenghao Peng, Quanyi Li, Ka Ming Hui, Chunxiao Liu, Bolei Zhou

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2021. Code and video can be found at: https://decisionforce.github.io/CoPO/

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.10190 2021-12-21 cs.AI 57%

Demanding and Designing Aligned Cognitive Architectures

Koen Holtman

专题命中 安全训练 :alignment(abstract);分类 cs.AI

Comments PERLS Workshop at 35th Conference on Neural Information Processing Systems (NeurIPS 2021). This arXiv version extends the workshop camera-ready version by adding four figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06185 2021-12-14 cs.RO cs.AI 57%

Multi-Agent Vulnerability Discovery for Autonomous Driving with Hazard Arbitration Reward

Weilin Liu, Ye Mu, Chao Yu, Xuefei Ning, Zhong Cao, Yi Wu, Shuang Liang, Huazhong Yang, Yu Wang

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.05434 2021-12-13 cs.AI 57%

A Reinforcement Learning-based Adaptive Control Model for Future Street Planning, An Algorithm and A Case Study

Qiming Ye, Yuxiang Feng, Jing Han, Marc Stettler, Panagiotis Angeloudis

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Proceeding for 57th ISOCARP World Planning Congress, Nov 8-11, 2021, Doha, Qatar

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06266 2021-12-08 cs.RO cs.LG cs.SY eess.SY 57%

Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning

Lukas Brunke, Melissa Greeff, Adam W. Hall, Zhaocong Yuan, Siqi Zhou, Jacopo Panerati, Angela P. Schoellig

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 36 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.05819 2021-11-17 cs.AI 57%

Look Before You Leap: Safe Model-Based Reinforcement Learning with Human Intervention

Yunkun Xu, Zhenyu Liu, Guifang Duan, Jiangcheng Zhu, Xiaolong Bai, Jianrong Tan

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments CoRL 2021 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.14870 2021-11-16 cs.AI 57%

A Scenario-Based Platform for Testing Autonomous Vehicle Behavior Prediction Models in Simulation

Francis Indaheng, Edward Kim, Kesav Viswanadha, Jay Shenoy, Jinkyu Kim, Daniel J. Fremont, Sanjit A. Seshia

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Accepted to the NeurIPS 2021 Workshop on Machine Learning for Autonomous Driving

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06831 2021-11-02 cs.AI cs.RO 57%

Safe Driving via Expert Guided Policy Optimization

Zhenghao Peng, Quanyi Li, Chunxiao Liu, Bolei Zhou

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02794 2021-10-26 stat.ML cs.LG 57%

Efficient Connected and Automated Driving System with Multi-agent Graph Reinforcement Learning

Tianyu Shi, Jiawei Wang, Yuankai Wu, Luis Miranda-Moreno, Lijun Sun

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.01296 2021-10-22 cs.LG 57%

A Safe Reinforcement Learning Architecture for Antenna Tilt Optimisation

Erik Aumayr, Saman Feghhi, Filippo Vannella, Ezeddin Al Hakim, Grigorios Iakovidis

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 6 pages, 3 figures, added copyright note

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02566 2021-10-07 eess.SY cs.LG cs.SY 57%

Adaptive control of a mechatronic system using constrained residual reinforcement learning

Tom Staessens, Tom Lefebvre, Guillaume Crevecoeur

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.00894 2021-10-05 cs.LG 57%

BRAC+: Improved Behavior Regularized Actor Critic for Offline Reinforcement Learning

Chi Zhang, Sanmukh Rao Kuppannagari, Viktor K Prasanna

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 16 pages. Accepted by ACML21

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.05997 2021-09-17 cs.LG cs.CR cs.LO cs.SC 57%

Verifying Quantized Neural Networks using SMT-Based Model Checking

Luiz Sena, Xidan Song, Erickson Alves, Iury Bessa, Edoardo Manino, Lucas Cordeiro, Eddie de Lima Filho

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Changes with respect to the previous version: improved explanation of our methodology in Section 3; improved and extended experimental evaluation in Section 4; added comparison with the state of the art in Section 4.5

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.05773 2021-08-30 cs.RO cs.AI cs.ET 57%

Approximate Computing for Robotic path planning -- Experimentation, Case Study and Practical Implications

Hrishav Bakul Barua

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Approximate Computing, Multi-robot Systems, Multi-agent Systems, Good Enough Computing, Green Computing, Robot Path Planning, Energy-Efficient Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.03952 2021-08-12 cs.LG cs.RO 57%

Safe Deep Reinforcement Learning for Multi-Agent Systems with Continuous Action Spaces

Ziyad Sheebaelhamd, Konstantinos Zisis, Athina Nisioti, Dimitris Gkouletsos, Dario Pavllo, Jonas Kohler

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments ICML 2021 Workshop on Reinforcement Learning for Real Life

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.10390 2021-07-23 cs.AI 57%

Reinforcement Learning Agent Training with Goals for Real World Tasks

Xuan Zhao, Marcos Campos

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Accepted to Reinforcement Learning for Real Life (RL4RealLife) Workshop in the 38th International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.09491 2021-07-23 cs.RO cs.AI 57%

Symbiotic System of Systems Design for Safe and Resilient Autonomous Robotics in Offshore Wind Farms

Daniel Mitchell, Jamie Blanche, Osama Zaki, Joshua Roe, Leo Kong, Samuel Harper, Valentin Robu, Theodore Lim, David Flynn

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments A preprint submit to IEEE Access Reliability Society Section

详情

展开后加载摘要…

URL PDF HTML 收藏