arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2211.02736 2024-04-11 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 62%

Discovering Closed-Loop Failures of Vision-Based Controllers via Reachability Analysis

Kaustav Chakraborty, Somil Bansal

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref IEEE Robotics and Automation Letters 8.5 (2023): 2692-2699

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12508 2024-04-05 cs.LG cs.AI 62%

SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation

Chongyu Fan, Jiancheng Liu, Yihua Zhang, Eric Wong, Dennis Wei, Sijia Liu

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by ICLR 2024 as a Spotlight paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16611 2024-04-03 cs.CL cs.AI cs.HC 62%

Understanding the Dataset Practitioners Behind Large Language Model Development

Crystal Qian, Emily Reif, Minsuk Kahng

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 2 figures. To be published in In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA '24). Revised to reflect updates from CHI LBW reviewer feedback

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00983 2024-04-02 cs.LG cs.AI 62%

Continual Learning for Smart City: A Survey

Li Yang, Zhipeng Luo, Shiming Zhang, Fei Teng, Tianrui Li

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Preprint. Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04357 2024-03-28 cs.CL cs.AI 62%

Dial-MAE: ConTextual Masked Auto-Encoder for Retrieval-based Dialogue Systems

Zhenpeng Su, Xing Wu, Wei Zhou, Guangyuan Ma, Songlin Hu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted by NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13257 2024-03-27 cs.CL cs.AI 62%

Visual Grounding Helps Learn Word Meanings in Low-Data Regimes

Chengxu Zhuang, Evelina Fedorenko, Jacob Andreas

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13830 2024-03-22 q-bio.BM cs.CL cs.LG 62%

Bridging Text and Molecule: A Survey on Multimodal Frameworks for Molecule

Yi Xiao, Xiangxin Zhou, Qiang Liu, Liang Wang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13514 2024-03-21 cs.CL cs.CY 62%

How Gender Interacts with Political Values: A Case Study on Czech BERT Models

Adnan Al Ali, Jindřich Libovický

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

Comments 11 pages, 2 figures; LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09795 2024-03-18 cs.CR cs.AI cs.CL 62%

Helpful or Harmful? Exploring the Efficacy of Large Language Models for Online Grooming Prevention

Ellie Prosser, Matthew Edwards

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07699 2024-03-15 cs.CV cs.AI cs.LG 62%

VeCLIP: Improving CLIP Training via Visual-enriched Captions

Zhengfeng Lai, Haotian Zhang, Bowen Zhang, Wentao Wu, Haoping Bai, Aleksei Timofeev, Xianzhi Du, Zhe Gan, Jiulong Shan, Chen-Nee Chuah, Yinfei Yang, Meng Cao

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments CV/ML

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06914 2024-03-13 cs.CL cs.AI 62%

MEND: Meta dEmonstratioN Distillation for Efficient and Effective In-Context Learning

Yichuan Li, Xiyao Ma, Sixing Lu, Kyumin Lee, Xiaohu Liu, Chenlei Guo

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12021 2024-03-12 cs.CL cs.AI 62%

Synergistic Anchored Contrastive Pre-training for Few-Shot Relation Extraction

Da Luo, Yanglei Gan, Rui Hou, Run Lin, Qiao Liu, Yuxiang Cai, Wannian Gao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12151 2024-03-05 cs.CL cs.AI 62%

Transformer-based Causal Language Models Perform Clustering

Xinbo Wu, Lav R. Varshney

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Added new experimental results and fixed some errors

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18807 2024-03-01 cs.CL cs.AI 62%

On the Decision-Making Abilities in Role-Playing using Large Language Models

Chenglei Shen, Guofu Xie, Xiao Zhang, Jun Xu

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18096 2024-02-29 cs.LG cs.AI 62%

No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

June Yong Yang, Byeongwook Kim, Jeongin Bae, Beomseok Kwon, Gunho Park, Eunho Yang, Se Jung Kwon, Dongsoo Lee

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16305 2024-02-27 cs.LG cs.AI 62%

Referee Can Play: An Alternative Approach to Conditional Generation via Model Inversion

Xuantong Liu, Tianyang Hu, Wenjia Wang, Kenji Kawaguchi, Yuan Yao

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09443 2024-02-16 eess.SP cs.AI cs.LG 62%

Review of algorithms for predicting fatigue using EEG

Ildar Rakhmatulin

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: text overlap with arXiv:2401.15766

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08088 2024-02-14 cs.AI cs.LG eess.IV 62%

Out-of-Distribution Detection and Data Drift Monitoring using Statistical Process Control

Ghada Zamzmi, Kesavan Venkatesh, Brandon Nelson, Smriti Prathapan, Paul H. Yi, Berkman Sahiner, Jana G. Delfino

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07344 2024-02-13 cs.LG cs.AI 62%

Measurement Scheduling for ICU Patients with Offline Reinforcement Learning

Zongliang Ji, Anna Goldenberg, Rahul G. Krishnan

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Extended Abstract presented at Machine Learning for Health (ML4H) symposium 2023, December 10th, 2023, New Orleans, United States, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06185 2024-02-12 cs.CV cs.AI cs.LG 62%

Development and validation of an artificial intelligence model to accurately predict spinopelvic parameters

Edward S. Harake, Joseph R. Linzey, Cheng Jiang, Rushikesh S. Joshi, Mark M. Zaki, Jaes C. Jones, Siri S. Khalsa, John H. Lee, Zachary Wilseck, Jacob R. Joseph, Todd C. Hollon, Paul Park

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 5 figures, to appear in Journal of Neurosurgery: Spine

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05627 2024-02-09 cs.LG cs.AI cs.CV q-bio.NC 62%

Binding Dynamics in Rotating Features

Sindy Löwe, Francesco Locatello, Max Welling

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04232 2024-02-08 cs.AI cs.CL 62%

Can Generative Agents Predict Emotion?

Ciaran Regan, Nanami Iwahashi, Shogo Tanaka, Mizuki Oka

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16123 2024-02-08 cs.HC cs.AI cs.CV cs.LG 62%

Looking for a better fit? An Incremental Learning Multimodal Object Referencing Framework adapting to Individual Drivers

Amr Gomaa, Guillermo Reyes, Michael Feld, Antonio Krüger

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in the Proceedings of the 29th International Conference on Intelligent User Interfaces (IUI'24), March 18--21, 2024, in Greenville, SC, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01516 2024-02-07 cs.CV cs.AI cs.LG cs.MM 62%

MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval

Zijun Long, George Killick, Richard McCreadie, Gerardo Aragon Camarasa

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03014 2024-02-06 cs.LG cs.AI 62%

Whom to Trust? Elective Learning for Distributed Gaussian Process Regression

Zewen Yang, Xiaobing Dai, Akshat Dubey, Sandra Hirche, Georges Hattab

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 9 pages, conference preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10246 2024-02-06 cs.LG cs.AI stat.ML 62%

Surprisal Driven $k$-NN for Robust and Interpretable Nonparametric Learning

Amartya Banerjee, Christopher J. Hazard, Jacob Beel, Cade Mack, Jack Xia, Michael Resnick, Will Goddin

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01091 2024-02-05 cs.CL cs.CY cs.SI 62%

Reading Between the Tweets: Deciphering Ideological Stances of Interconnected Mixed-Ideology Communities

Zihao He, Ashwin Rao, Siyi Guo, Negar Mokhberian, Kristina Lerman

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.12241 2024-02-05 cs.CY cs.AI 62%

Positive AI: Key Challenges in Designing Artificial Intelligence for Wellbeing

Willem van der Maden, Derek Lomas, Malak Sadek, Paul Hekkert

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03674 2024-02-05 cs.LG cs.AI cs.SE 62%

Machine Learning with Requirements: a Manifesto

Eleonora Giunchiglia, Fergus Imrie, Mihaela van der Schaar, Thomas Lukasiewicz

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01053 2024-02-05 cs.CL cs.AI 62%

Plan-Grounded Large Language Models for Dual Goal Conversational Settings

Diogo Glória-Silva, Rafael Ferreira, Diogo Tavares, David Semedo, João Magalhães

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏