arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7945 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7945 篇

2501.14294 2025-03-04 cs.CL cs.AI 81%

Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes

Sullam Jeoung, Yubin Ge, Haohan Wang, Jana Diesner

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18685 2025-02-27 cs.AI cs.CL cs.HC 81%

Speaking the Right Language: The Impact of Expertise Alignment in User-AI Interactions

Shramay Palta, Nirupama Chandrasekaran, Rachel Rudinger, Scott Counts

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments arXiv Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14047 2025-02-21 cs.LG cs.AI stat.ML 81%

Towards a Learning Theory of Representation Alignment

Francesco Insulla, Shuo Huang, Lorenzo Rosasco

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12552 2025-02-19 cs.CY cs.AI 81%

LLM Safety for Children

Prasanjit Rath, Hari Shrawgi, Parag Agrawal, Sandipan Dandapat

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06834 2025-02-18 cs.MA cs.AI cs.CY 81%

Investigating social alignment via mirroring in a system of interacting language models

Harvey McGuinness, Tianyu Wang, Carey E. Priebe, Hayden Helm

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07803 2025-02-13 cs.AI cs.LG 81%

Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment

Cheryl Li, Tianyuan Xu, Yiwen Guo

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06827 2025-02-12 cs.LG cs.AI cs.GR 81%

Learning to Synthesize Compatible Fashion Items Using Semantic Alignment and Collocation Classification: An Outfit Generation Framework

Dongliang Zhou, Haijun Zhang, Kai Yang, Linlin Liu, Han Yan, Xiaofei Xu, Zhao Zhang, Shuicheng Yan

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments This paper was accepted by IEEE TNNLS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19309 2025-02-03 cs.LG cs.CL 81%

Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment

Gregor Bachmann, Sotiris Anagnostidis, Albert Pumarola, Markos Georgopoulos, Artsiom Sanakoyeu, Yuming Du, Edgar Schönfeld, Ali Thabet, Jonas Kohler

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19087 2025-01-28 cs.CV cs.AI cs.LG q-bio.QM 81%

Dimensions underlying the representational alignment of deep neural networks with humans

Florian P. Mahner, Lukas Muttenthaler, Umut Güçlü, Martin N. Hebart

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03263 2025-01-24 cs.LG cs.AI 81%

Test-time Adaptation for Regression by Subspace Alignment

Kazuki Adachi, Shin'ya Yamaguchi, Atsutoshi Kumagai, Tomoki Hamagami

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06490 2025-01-23 cs.CL cs.LG 81%

Sequential Classification of Aviation Safety Occurrences with Natural Language Processing

Aziida Nanyonga, Hassan Wasswa, Ugur Turhan, Oleksandra Molloy, Graham Wild

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.LG

Journal ref AIAA AVIATION 2023 Forum (p. 4325)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09685 2025-01-22 cs.AI cs.LG q-bio.QM stat.ML 81%

Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review

Masatoshi Uehara, Yulai Zhao, Chenyu Wang, Xiner Li, Aviv Regev, Sergey Levine, Tommaso Biancalani

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments We plan to add more content and codes. Please let us know if there are any comments or missing citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02186 2025-01-16 cs.CV cs.AI cs.LG 81%

Identifying Spurious Correlations using Counterfactual Alignment

Joseph Paul Cohen, Louis Blankemeier, Akshay Chaudhari

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Accepted to Transactions on Machine Learning Research (TMLR), Code: https://github.com/ieee8023/latentshift

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06564 2025-01-14 cs.CL cs.LG 81%

Natural Language Processing and Deep Learning Models to Classify Phase of Flight in Aviation Safety Occurrences

Aziida Nanyonga, Hassan Wasswa, Oleksandra Molloy, Ugur Turhan, Graham Wild

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.LG

Comments NLP, Aviation reports, Text analysis, Deep learning algorithms, Flight phase classification

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06210 2025-01-14 cs.CL cs.LG 81%

Applications of natural language processing in aviation safety: A review and qualitative analysis

Aziida Nanyonga, Keith Joiner, Ugur Turhan, Graham Wild

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03681 2025-01-08 cs.CL cs.AI 81%

SLAM: Towards Efficient Multilingual Reasoning via Selective Language Alignment

Yuchun Fan, Yongyu Mu, Yilin Wang, Lei Huang, Junhao Ruan, Bei Li, Tong Xiao, Shujian Huang, Xiaocheng Feng, Jingbo Zhu

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by COLING 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13771 2024-12-19 cs.IR cs.AI cs.CL 81%

Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic Tokenization

Guanghan Li, Xun Zhang, Yufei Zhang, Yifan Yin, Guojun Yin, Wei Lin

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments 7 pages, 3 figures, AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06527 2024-12-18 cs.CL cs.AI 81%

Rethinking the Alignment of Psychotherapy Dialogue Generation with Motivational Interviewing Strategies

Xin Sun, Xiao Tang, Abdallah El Ali, Zhuying Li, Pengjie Ren, Jan de Wit, Jiahuan Pei, Jos A. Bosch

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19385 2024-12-02 cs.LG cs.AI eess.SP 81%

Zero-Forget Preservation of Semantic Communication Alignment in Distributed AI Networks

Jingzhi Hu, Geoffrey Ye Li

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16824 2024-11-27 cs.CV cs.AI cs.LG 81%

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge

Yaqi Zhao, Yuanyang Yin, Lin Li, Mingan Lin, Victor Shea-Jay Huang, Siwei Chen, Weipeng Chen, Baoqun Yin, Zenan Zhou, Wentao Zhang

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13018 2024-11-27 q-bio.NC cs.AI cs.LG cs.NE 81%

Getting aligned on representational alignment

Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller, Andi Peng, Andreea Bobu, Been Kim, Bradley C. Love, Christopher J. Cueva, Erin Grant, Iris Groen, Jascha Achterberg, Joshua B. Tenenbaum, Katherine M. Collins, Katherine L. Hermann, Kerem Oktar, Klaus Greff, Martin N. Hebart, Nathan Cloos, Nikolaus Kriegeskorte, Nori Jacoby, Qiuyi Zhang, Raja Marjieh, Robert Geirhos, Sherol Chen, Simon Kornblith, Sunayana Rane, Talia Konkle, Thomas P. O'Connell, Thomas Unterthiner, Andrew K. Lampinen, Klaus-Robert Müller, Mariya Toneva, Thomas L. Griffiths

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 51 pages; Working paper (changes to be made in upcoming revisions)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16442 2024-11-26 cs.LG cs.AI 81%

TIFeD: a Tiny Integer-based Federated learning algorithm with Direct feedback alignment

Luca Colombo, Alessandro Falcetta, Manuel Roveri

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Journal ref Proceedings of the Third International Conference on AI-ML Systems, 2023, pp. 1-8

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14129 2024-11-26 cs.CL cs.AI cs.CV 81%

AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability

Fei Zhao, Taotian Pang, Chunhui Li, Zhen Wu, Junjie Guo, Shangyu Xing, Xinyu Dai

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18634 2024-11-19 cs.LG cs.CL stat.ML 81%

A Theoretical Understanding of Self-Correction through In-context Alignment

Yifei Wang, Yuyang Wu, Zeming Wei, Stefanie Jegelka, Yisen Wang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments Accepted at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07013 2024-11-12 cs.LG cs.AI cs.NI 81%

A neural-network based anomaly detection system and a safety protocol to protect vehicular network

Marco Franceschini

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments Master's thesis 2023-2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.24198 2024-11-04 cs.CL cs.LG cs.SE 81%

SelfCodeAlign: Self-Alignment for Code Generation

Yuxiang Wei, Federico Cassano, Jiawei Liu, Yifeng Ding, Naman Jain, Zachary Mueller, Harm de Vries, Leandro von Werra, Arjun Guha, Lingming Zhang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12298 2024-10-18 cs.CL cs.AI 81%

Pyramid-Driven Alignment: Pyramid Principle Guided Integration of Large Language Models and Knowledge Graphs

Lei Sun, Xinchen Wang, Youdi Li

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03360 2024-10-15 cs.LG cs.AI 81%

Prioritize Alignment in Dataset Distillation

Zekai Li, Ziyao Guo, Wangbo Zhao, Tianle Zhang, Zhi-Qi Cheng, Samir Khaki, Kaipeng Zhang, Ahmad Sajedi, Konstantinos N Plataniotis, Kai Wang, Yang You

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 19 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15048 2024-10-11 cs.CL cs.AI 81%

Unlocking the Power of Large Language Models for Entity Alignment

Xuhui Jiang, Yinghan Shen, Zhichao Shi, Chengjin Xu, Wei Li, Zixuan Li, Jian Guo, Huawei Shen, Yuanzhuo Wang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13642 2024-10-10 cs.LG cs.AI cs.SY eess.SY 81%

Safety Margins for Reinforcement Learning

Alexander Grushin, Walt Woods, Alvaro Velasquez, Simon Khan

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 2 pages, 2 figures. Presented at the 2023 IEEE Conference on Artificial Intelligence (CAI), Santa Clara, CA

详情

展开后加载摘要…

URL PDF HTML 收藏