arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3281 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3281 篇

2407.08964 2024-07-15 cs.LG cs.RO 57%

Communication-Aware Reinforcement Learning for Cooperative Adaptive Cruise Control

Sicong Jiang, Seongjin Choi, Lijun Sun

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08735 2024-07-12 cs.RO cs.AI cs.SY eess.SY 57%

Real-Time Anomaly Detection and Reactive Planning with Large Language Models

Rohan Sinha, Amine Elhafsi, Christopher Agia, Matthew Foutter, Edward Schmerling, Marco Pavone

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Accepted to Robotics: Science and Systems (RSS) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05557 2024-07-09 cs.AI 57%

$R^2$-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning

Mintong Kang, Bo Li

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10329 2024-07-04 stat.AP cs.AI 57%

Causal inference approach to appraise long-term effects of maintenance policy on functional performance of asphalt pavements

Lingyun You, Nanning Guo, Zhengwu Long, Fusong Wang, Chundi Si, Aboelkasim Diab

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments The arXiv version needs to be withdrawn since the model needs to be validated and updated with advanced machine learning technologies to enhance the accuracy of the model, and there are some crucial definition errors of symbols in the arXiv version

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02245 2024-07-03 cs.RO cs.AI 57%

Safe CoR: A Dual-Expert Approach to Integrating Imitation Learning and Safe Reinforcement Learning Using Constraint Rewards

Hyeokjin Kwon, Gunmin Lee, Junseo Lee, Songhwai Oh

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Accepted to the Proc. of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01850 2024-07-03 cs.CL 57%

Purple-teaming LLMs with Adversarial Defender Training

Jingyan Zhou, Kun Li, Junan Li, Jiawen Kang, Minda Hu, Xixin Wu, Helen Meng

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15176 2024-07-03 cs.CL cs.SI 57%

Robust Stance Detection: Understanding Public Perceptions in Social Media

Nayoung Kim, David Mosallanezhad, Lu Cheng, Michelle V. Mancenido, Huan Liu

专题命中 安全训练 :safety(abstract);分类 cs.CL

Journal ref ASONAM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13025 2024-06-21 cs.LG cs.RO cs.SY eess.SY 57%

ABNet: Attention BarrierNet for Safe and Scalable Robot Learning

Wei Xiao, Tsun-Hsuan Wang, Daniela Rus

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12298 2024-06-19 cs.AI physics.ao-ph 57%

Research on Dangerous Flight Weather Prediction based on Machine Learning

Haoxing Liu, Renjie Xie, Haoshen Qin, Yizhou Li

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12117 2024-06-19 cs.CL 57%

Decoding the Narratives: Analyzing Personal Drug Experiences Shared on Reddit

Layla Bouzoubaa, Elham Aghakhani, Max Song, Minh Trinh, Rezvaneh Rezapour

专题命中 安全训练 :safety(abstract);分类 cs.CL

Comments Findings of the Association for Computational Linguistics: ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12594 2024-06-19 cs.LG cs.RO 57%

State-wise Constrained Policy Optimization

Weiye Zhao, Rui Chen, Yifan Sun, Tianhao Wei, Changliu Liu

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Published in Transactions of Machine Learning Research

Journal ref Transactions on Machine Learning Research (2024). ISSN: 2835-8856. https://openreview.net/forum?id=NgK5etmhz9

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11036 2024-06-18 cs.CL cs.CR 57%

garak: A Framework for Security Probing Large Language Models

Leon Derczynski, Erick Galinkin, Jeffrey Martin, Subho Majumdar, Nanna Inie

专题命中 安全训练 :alignment(abstract);分类 cs.CL

Comments https://garak.ai

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08884 2024-06-14 cs.CV cs.LG stat.ML 57%

The Penalized Inverse Probability Measure for Conformal Classification

Paul Melki, Lionel Bombrun, Boubacar Diallo, Jérôme Dias, Jean-Pierre da Costa

专题命中 安全训练 :trustworthy(abstract);分类 cs.LG

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE/CVF, Jun 2024, Seattle, United States

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08822 2024-06-14 cs.CV cs.AI 57%

Computer vision-based model for detecting turning lane features on Florida's public roadways

Richard Boadu Antwi, Samuel Takyi, Kimollo Michael, Alican Karaer, Eren Erman Ozguven, Ren Moses, Maxim A. Dulebenets, Thobias Sando

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00282 2024-06-03 cs.LG 57%

Conflict-Averse Gradient Aggregation for Constrained Multi-Objective Reinforcement Learning

Dohyeong Kim, Mineui Hong, Jeongho Park, Songhwai Oh

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18209 2024-05-29 cs.RO cs.LG 57%

Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving

Zhi Zheng, Shangding Gu

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04117 2024-05-22 cs.LG cs.SY eess.SY 57%

Ablation Study of How Run Time Assurance Impacts the Training and Performance of Reinforcement Learning Agents

Nathaniel Hamilton, Kyle Dunlap, Taylor T Johnson, Kerianne L Hobbs

专题命中 安全训练 :safety(abstract);分类 cs.LG

Journal ref 2023 IEEE 9th International Conference on Space Mission Challenges for Information Technology (SMC-IT), 2023, pp. 45-55

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11030 2024-05-21 cs.CL 57%

The Unappreciated Role of Intent in Algorithmic Moderation of Social Media Content

Xinyu Wang, Sai Koneru, Pranav Narayanan Venkit, Brett Frischmann, Sarah Rajtmajer

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05146 2024-05-10 cs.AI 57%

Hybrid Convolutional Neural Networks with Reliability Guarantee

Hans Dermot Doran, Suzana Veljanovska

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 2024 54th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN 2024). Dependable and Secure Machine Learning Workshop (DSML 2024), Brisbane, Australia, June 24-27, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04015 2024-05-08 cs.AI cs.LO 57%

Certified Policy Verification and Synthesis for MDPs under Distributional Reach-avoidance Properties

S. Akshay, Krishnendu Chatterjee, Tobias Meggendorfer, Đorđe Žikelić

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Extended version of a paper accepted at IJCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03023 2024-04-05 cs.HC cs.AI 57%

Toward Safe Evolution of Artificial Intelligence (AI) based Conversational Agents to Support Adolescent Mental and Sexual Health Knowledge Discovery

Jinkyung Park, Vivek Singh, Pamela Wisniewski

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments This paper has been peer-reviewed and presented at the "CHI 2024 Workshop on Child-centred AI Design, May 11, 2024, Honolulu, HI, USA."

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03773 2024-04-01 cs.RO cs.AI 57%

Safe Explicable Planning

Akkamahadevi Hanni, Andrew Boateng, Yu Zhang

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.01563 2024-04-01 cs.RO cs.LG 57%

Data-efficient, Explainable and Safe Box Manipulation: Illustrating the Advantages of Physical Priors in Model-Predictive Control

Achkan Salehi, Stephane Doncieux

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments accepted for publication by l4dc 2024, 12 pages (with references), 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09749 2024-03-22 cs.CL 57%

Facilitating NSFW Text Detection in Open-Domain Dialogue Systems via Knowledge Distillation

Huachuan Qiu, Shuai Zhang, Hongliang He, Anqi Li, Zhenzhong Lan

专题命中 安全训练 :safety(abstract);分类 cs.CL

Comments As we have submitted a final version arXiv:2403.13250, we decide to withdraw it

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09509 2024-03-20 cs.SE cs.LG 57%

On STPA for Distributed Development of Safe Autonomous Driving: An Interview Study

Ali Nouri, Christian Berger, Fredrik Törner

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted at SEAA. 8 pages, 2 figures

Journal ref 2023 49th Euromicro Conference on Software Engineering and Advanced Applications (SEAA), Durres, Albania, 2023, pp. 5-12

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03173 2024-03-20 cs.CL 57%

$\mathcal{B}$-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis

Zishun Yu, Yunzhe Tao, Liyu Chen, Tao Sun, Hongxia Yang

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19282 2024-03-19 cs.CL 57%

WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset

Jiantao Qiu, Haijun Lv, Zhenjiang Jin, Rui Wang, Wenchang Ning, Jia Yu, ChaoBin Zhang, Zhenxiang Li, Pei Chu, Yuan Qu, Jin Shi, Lindong Lu, Runyu Peng, Zhiyuan Zeng, Huanze Tang, Zhikai Lei, Jiawei Hong, Keyu Chen, Zhaoye Fei, Ruiliang Xu, Wei Li, Zhongying Tu, Lin Dahua, Yu Qiao, Hang Yan, Conghui He

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03351 2024-03-19 cs.LG cs.RO 57%

Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization

Kun Lei, Zhengmao He, Chenhao Lu, Kaizhe Hu, Yang Gao, Huazhe Xu

专题命中 安全训练 :alignment(abstract);分类 cs.LG

Comments Our website: https://lei-kun.github.io/uni-o4/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12359 2024-03-19 cs.MA cs.LG 57%

MARVEL: Multi-Agent Reinforcement-Learning for Large-Scale Variable Speed Limits

Yuhang Zhang, Marcos Quinones-Grueiro, Zhiyao Zhang, Yanbing Wang, William Barbour, Gautam Biswas, Daniel Work

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05693 2024-03-15 cs.LG 57%

Shielded Deep Reinforcement Learning for Complex Spacecraft Tasking

Robert Reed, Hanspeter Schaub, Morteza Lahijanian

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 9 pages, 2 figures, 2 tables, ACC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏