arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7968 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7968 篇

2406.13357 2024-06-21 cs.CL cs.SD eess.AS 79%

Transferable speech-to-text large language model alignment module

Boyong Wu, Chao Yan, Haoran Pu

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted by InterSpeech 2024; 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05678 2024-06-21 cs.HC cs.CL 79%

Beyond Prompts: Learning from Human Communication for Enhanced AI Intent Alignment

Yoonsu Kim, Kihoon Son, Seoyoung Kim, Juho Kim

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06228 2024-06-12 cs.CL 79%

Understanding Cross-Lingual Alignment -- A Survey

Katharina Hämmerl, Jindřich Libovický, Alexander Fraser

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Camera-ready version, ACL Findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06049 2024-06-11 cs.CY 79%

Enhancing Food Safety in Supply Chains: The Potential Role of Large Language Models in Preventing Campylobacter Contamination

Asaf Tzachor

专题命中 其他安全 :safety(title,abstract);分类 cs.CY

Comments 29 pages, 1 figure, 3 boxes

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05505 2024-06-11 cs.IR cs.AI 79%

I-SIRch: AI-Powered Concept Annotation Tool For Equitable Extraction And Analysis Of Safety Insights From Maternity Investigations

Mohit Kumar Singh, Georgina Cosma, Patrick Waterson, Jonathan Back, Gyuchan Thomas Jun

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04799 2024-06-10 cs.CL 79%

Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages

Shih-Cheng Huang, Pin-Zu Li, Yu-Chi Hsu, Kuang-Ming Chen, Yu Tung Lin, Shih-Kai Hsiao, Richard Tzong-Han Tsai, Hung-yi Lee

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments ACL 2024 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02915 2024-06-06 cs.CV cs.LG 79%

Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models

Jinhao Li, Haopeng Li, Sarah Erfani, Lei Feng, James Bailey, Feng Liu

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments 22 pages, 16 figures, published to ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00117 2024-05-29 cs.CL 79%

BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B

Pranav Gade, Simon Lermen, Charlie Rogers-Smith, Jeffrey Ladish

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00916 2024-05-29 cs.CL cs.SD eess.AS 79%

BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Chen Wang, Minpeng Liao, Zhongqiang Huang, Jinliang Lu, Junhong Wu, Yuchen Liu, Chengqing Zong, Jiajun Zhang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15787 2024-05-28 cs.IR cs.CL 79%

Extracting chemical food safety hazards from the scientific literature automatically using large language models

Neris Özen, Wenjuan Mu, Esther D. van Asselt, Leonieke M. van den Bulk

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

Comments 31 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15430 2024-05-27 cs.LG cs.LO 79%

Counterexample-Guided Repair of Reinforcement Learning Systems Using Safety Critics

David Boetius, Stefan Leue

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments 7 pages + references

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15202 2024-05-27 cs.CL cs.CR 79%

Cross-Task Defense: Instruction-Tuning LLMs for Content Safety

Yu Fu, Wen Xiao, Jia Chen, Jiachen Li, Evangelos Papalexakis, Aichi Chien, Yue Dong

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

Comments accepted to NAACL2024 TrustNLP workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07162 2024-05-17 cs.RO cs.AI 79%

Learning Reward for Robot Skills Using Large Language Models via Self-Alignment

Yuwei Zeng, Yao Mu, Lin Shao

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08154 2024-05-15 cs.HC cs.AI 79%

LLM Theory of Mind and Alignment: Opportunities and Risks

Winnie Street

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Journal ref Proceedings of Workshop on Theory of Mind in Human-AI Interaction at CHI 2024 (ToMinHAI at CHI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03939 2024-05-08 cs.CL 79%

Long Context Alignment with Short Instructions and Synthesized Positions

Wenhao Wu, Yizhong Wang, Yao Fu, Xiang Yue, Dawei Zhu, Sujian Li

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments preview

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03699 2024-05-08 cs.HC cs.AI 79%

HCC Is All You Need: Alignment-The Sensible Kind Anyway-Is Just Human-Centered Computing

Eric Gilbert

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00462 2024-05-06 cs.LG cs.RO 79%

Zero-shot Safety Prediction for Autonomous Robots with Foundation World Models

Zhenjiang Mao, Siqi Dai, Yuang Geng, Ivan Ruchkin

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments Presented at the Back to the Future-Robot Learning Going Probabilistic Workshop, co-located with ICRA 2024. https://openreview.net/forum?id=gHhBNIq9Cs

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08774 2024-04-30 physics.soc-ph cs.AI 79%

Discussion of Loop Expansion and Introduction of Series Cutting Functions to Local Potential Approximation: Complexity Analysis Using Green's Functions, Cutting Of Nth-Order Social Interactions For Progressive Safety

Yasuko Kawahata

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments In this study, we focus on the aforementioned paper, "Examination Kubo-Matsubara Green's Function Of The Edwards-Anderson Model: Extreme Value Information Flow Of Nth-Order Interpolated Extrapolation Of Zero Phenomena Using The Replica Method (2024)"

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05183 2024-04-09 cs.CV cs.LG 79%

Progressive Alignment with VLM-LLM Feature to Augment Defect Classification for the ASE Dataset

Chih-Chung Hsu, Chia-Ming Lee, Chun-Hung Sun, Kuang-Ming Wu

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments MULA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04659 2024-04-09 cs.CL 79%

Multilingual Pretraining and Instruction Tuning Improve Cross-Lingual Knowledge Alignment, But Only Shallowly

Changjiang Gao, Hongda Hu, Peng Hu, Jiajun Chen, Jixing Li, Shujian Huang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01632 2024-04-03 cs.LG cs.SY eess.SY 79%

Enhancing Functional Safety in Automotive AMS Circuits through Unsupervised Machine Learning

Ayush Arunachalam, Ian Kintz, Suvadeep Banerjee, Arnab Raha, Xiankun Jin, Fei Su, Viswanathan Pillai Prasanth, Rubin A. Parekhji, Suriyaprakash Natarajan, Kanad Basu

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments 12 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18435 2024-03-28 cs.IR cs.CL 79%

DELTA: Pre-train a Discriminative Encoder for Legal Case Retrieval via Structural Word Alignment

Haitao Li, Qingyao Ai, Xinyan Han, Jia Chen, Qian Dong, Yiqun Liu, Chong Chen, Qi Tian

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14149 2024-03-27 cs.CV cs.AI 79%

TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification

Qinying Liu, Wei Wu, Kecheng Zheng, Zhan Tong, Jiawei Liu, Yu Liu, Wei Chen, Zilei Wang, Yujun Shen

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16904 2024-03-26 cs.AI cs.CR 79%

Multi-Agent Optimization for Safety Analysis of Cyber-Physical Systems: Position Paper

Önder Gürcan, Nataliya Yakymets, Sara Tucci-Piergiovanni, Ansgar Radermacher

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments 13 pages, 2 figures, 1 table, "2nd International Workshop on Emerging Ideas and Trends in Engineering of Cyber-Physical Systems, part of Cyber-Physical Systems Week, April 2015, Seattle, USA"

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16289 2024-03-26 cs.AI 79%

Engineering Safety Requirements for Autonomous Driving with Large Language Models

Ali Nouri, Beatriz Cabrero-Daniel, Fredrik Törner, Hȧkan Sivencrona, Christian Berger

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments Accepted in 32nd IEEE International Requirements Engineering 2024 conference, Iceland

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14203 2024-03-22 cs.CV cs.AI 79%

Unsupervised Audio-Visual Segmentation with Modality Alignment

Swapnil Bhosale, Haosen Yang, Diptesh Kanojia, Jiangkang Deng, Xiatian Zhu

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04527 2024-03-20 cs.IR cs.AI 79%

RA-Rec: An Efficient ID Representation Alignment Framework for LLM-based Recommendation

Xiaohan Yu, Li Zhang, Xin Zhao, Yue Wang, Zhongrui Ma

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11368 2024-03-19 cs.RO cs.AI 79%

Driving Style Alignment for LLM-powered Driver Agent

Ruoxuan Yang, Xinyue Zhang, Anais Fernandez-Laaksonen, Xin Ding, Jiangtao Gong

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06888 2024-03-14 physics.data-an cs.LG physics.app-ph 79%

Process signature-driven high spatio-temporal resolution alignment of multimodal data

Abhishek Hanchate, Himanshu Balhara, Vishal S. Chindepalli, Satish T. S. Bukkapatnam

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06259 2024-03-13 cs.CL 79%

Self-Alignment with Instruction Backtranslation

Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason Weston, Mike Lewis

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments ICLR2024 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏