arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7968 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7968 篇

2211.16550 2023-05-29 cs.CL cs.AI cs.NE 76%

Soft Alignment Objectives for Robust Adaptation of Language Generation

Michal Štefánik, Marek Kadlčík, Petr Sojka

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

Comments Annual Meeting of The ACL 2023: Main conference long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06386 2023-05-12 cs.CV cs.AI cs.HC cs.LG 76%

Text-To-Concept (and Back) via Cross-Model Alignment

Mazda Moayeri, Keivan Rezaei, Maziar Sanjabi, Soheil Feizi

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments Accepted to ICML 2023 and CVPR4XAI workshop 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.15259 2023-05-10 cs.LG cs.AI cs.CR 76%

On the Impossible Safety of Large AI Models

El-Mahdi El-Mhamdi, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Lê-Nguyên Hoang, Rafael Pinot, Sébastien Rouault, John Stephan

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Comments 40 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00902 2023-02-06 cs.LG cs.CL cs.CV 76%

Language Quantized AutoEncoders: Towards Unsupervised Text-Image Alignment

Hao Liu, Wilson Yan, Pieter Abbeel

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

Comments Fixed typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03525 2022-04-08 cs.LG cs.AI 76%

Temporal Alignment for History Representation in Reinforcement Learning

Aleksandr Ermolov, Enver Sangineto, Nicu Sebe

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments ICPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13411 2022-03-28 cs.RO cs.AI cs.LG cs.SY eess.SY 76%

Reshaping Robot Trajectories Using Natural Language Commands: A Study of Multi-Modal Data Alignment Using Transformers

Arthur Bucker, Luis Figueredo, Sami Haddadin, Ashish Kapoor, Shuang Ma, Rogerio Bonatti

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.14659 2021-03-30 cs.AI cs.LG 76%

Alignment of Language Agents

Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, Geoffrey Irving

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.06907 2019-10-16 cs.LG cs.AI math.OC 76%

Techniques for Adversarial Examples Threatening the Safety of Artificial Intelligence Based Systems

Utku Kose

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Comments International Science and Innovation Congress 2019, pp. 643-655, 13 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.01267 2018-11-06 cs.AI cs.CY cs.HC 76%

Legible Normativity for AI Alignment: The Value of Silly Rules

Dylan Hadfield-Menell, McKane Andrus, Gillian K. Hadfield

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22226 2026-06-23 cs.GT cs.AI cs.IT math.IT 新提交 76%

Quantifying Theoretical AI Alignment Guarantees: Receiver-Utility Bounds in Bayesian Persuasion

量化理论上的AI对齐保证:贝叶斯说服中的接收者效用界

Eric Yachbes, Eva Tardos

机构 * Cornell University(康奈尔大学)

专题命中 其他安全 :alignment(title,comments);分类 cs.AI

AI总结 通过贝叶斯说服模型,研究AI发送者优化错位目标时,人类接收者仍能获得多少有用信息,证明接收者效用比不超过3/2,并给出紧性下界。

Comments 12 pages, EC 2026 Poster and EC 2026 Incentive-Based AI Alignment Workshop Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19586 2025-10-28 cs.CL q-bio.NC 76%

Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment

Hanlin Wu, Xufeng Duan, Zhenguang Cai

机构 * Department of Linguistics and Modern Languages, The Chinese University of Hong Kong(语言学与现代语言系,香港中文大学) Brain and Mind Institute, The Chinese University of Hong Kong(脑与心智研究所,香港中文大学)

专题命中 其他安全 :alignment(title,comments);分类 cs.CL

Comments Hanlin Wu, Xufeng Duan, and Zhenguang Cai. 2025. Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 135-143, Albuquerque, New Mexico, USA. Association for Computational Linguistics. https://aclanthology.org/2025.cmcl-1.18/

Journal ref In Proceedings of CMCL, pages 135-143, ACL (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22925 2026-07-28 cs.CL cs.AI cs.LG 新提交 75%

Not All LLM Reasoning is Visible in the Chain-of-Thought

并非所有大语言模型的推理都能在思维链中体现

Vatsal Baherwani, Tom Goldstein, Ashwinee Panda

机构 * New York University(纽约大学) University of Maryland(马里兰大学) TogetherAI

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究探讨大语言模型输出令牌是否体现所有推理,发现前沿模型存在利用无关填充令牌提升合成推理任务性能的不可见推理现象,评估多个模型,揭示填充令牌益处因模型和令牌而异,还表明其能服务隐藏目标,且强化学习等方法无法使填充令牌益处在测试时持续。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04885 2026-06-15 cs.CL cs.AI cs.LG 版本更新 75%

CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters

CuMA: 通过人口统计感知的适配器混合使大语言模型与稀疏文化价值观对齐

Ao Sun, Xiaoyu Wang, Zhe Tan, Yu Li, Jiachen Zhu, Yuheng Jia, Shu Su

机构 * Southeast University(东南大学) ByteDance Inc.(字节跳动公司) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用重点实验室(东南大学),中华人民共和国教育部,中国)

专题命中 其他安全 :alignment(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 提出CuMA框架,通过人口统计感知路由将冲突梯度分离到专家子空间,解决密集模型在多文化对齐中的均值崩溃问题,在WorldValuesBench等基准上取得最优性能。

Comments ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22005 2026-05-26 cs.LG cs.AI cs.CL 75%

Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have)

检查你的大语言模型的秘密词典!五行代码揭示你的大语言模型学到了什么(包括它不应该学到的)

Hisashi Miyashita

机构 * Mgnite Inc.(Mgnite公司)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过对lm_head权重矩阵进行奇异值分解(仅需五行PyTorch代码且无需模型推理),直接从模型权重中揭示可解释的语义子空间,并发现模型训练数据组成和策展哲学。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02064 2026-05-19 cs.LG cs.AI cs.CL 75%

Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct

对Llama3-8b-Instruct自生成文本识别能力的检查与控制

Christopher Ackerman, Nina Panickssery

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究探讨了LLM是否能识别自身生成的文本,发现Llama3-8b-Instruct模型能够区分自身输出与人类输出,并通过残差流中的特定向量控制其行为和感知,揭示了模型自我归属的认知机制。

Comments 10 pages, 13 figs, 2 tables, accepted as conference paper to ICLR 2025

Journal ref The Thirteenth International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20164 2026-05-12 cs.LG cs.AI cs.CL 75%

What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering

计划是什么?LLMs中隐式规划的度量及其在押韵生成和问答中的应用

Jim Maar, Denis Paperno, Callum Stuart McDougall, Neel Nanda

机构 * HPI / University of Potsdam(HPI/波茨坦大学) Utrecht University(乌特勒支大学) Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出简单方法评估LLM隐式规划,通过押韵生成和问答案例展示其可扩展性,发现隐式规划在1B参数模型中普遍存在,为AI安全提供新视角。

Comments 41 pages, 34 figures, Accepted at ICLR 2026, Code available at https://github.com/Jim-Maar/implicit-planning-in-llms

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17614 2026-04-21 cs.AI cs.CL cs.LG 75%

Characterizing Model-Native Skills

刻画模型内禀技能

Feiyang Kang, Mahavir Dabas, Myeongseob Ko, Ruoxi Jia

机构 * Virginia Tech(弗吉尼亚理工学院)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出模型内禀技能的刻画方法,通过从序列激活中恢复紧凑正交基,实现行为变化轴的自组织,验证了在推理和安全对齐中的有效性,优于人类定义的技能。

Comments We argue that when the goal is to intervene on model behavior, skill characterization should be *model-native*: grounded in the model's own representations rather than imposed through external ontologies

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03230 2026-03-04 cs.LG cs.AI cs.CL math.OC 75%

DiaBlo: Diagonal Blocks Are Sufficient For Finetuning

DiaBlo: 对角块足以用于微调

Selcuk Gurses, Aozhong Zhang, Yanxia Deng, Xun Dong, Xin Li, Naigang Wang, Penghang Yin, Zi Yang

机构 * University at Albany, SUNY(纽约州立大学阿尔巴尼分校) IBM T. J. Watson Research Center(IBM 汤普逊·杰·沃森研究中心)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 DiaBlo是一种仅更新模型权重矩阵对角块的参数高效微调方法,通过消除低秩矩阵乘积需求,实现稳定收敛和高效训练。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24940 2026-01-01 cs.AI cs.CL cs.LG 75%

Iterative Deployment Improves Planning Skills in LLMs

迭代部署提升大语言模型的规划能力

Augusto B. Corrêa, Yoav Gelberg, Luckeciano C. Melo, Ilia Shumailov, André G. Pereira, Yarin Gal

机构 * University of Oxford(牛津大学) AI Sequrity Company(AI安全公司) UFRGS(乌拉圭联邦大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过迭代部署大语言模型,利用用户编纂的数据提升规划能力,展现隐含奖励函数的强化学习机制,具有AI安全和训练制度替代的双重意义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11402 2025-10-02 cs.LG cs.AI cs.CL 75%

LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model

Marcel Mateos Salles, Praney Goyal, Pradyut Sekhsaria, Hai Huang, Randall Balestriero

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 46 pages, 17 figures, 26 tables. Submitted for publication. for associated blog post, see https://pradyut3501.github.io/lora-spur-corr/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07448 2025-08-05 cs.LG cs.AI cs.CL 75%

LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation

Juzheng Zhang, Jiacheng You, Ashwinee Panda, Tom Goldstein

机构 * University of Maryland(马里兰大学) Tsinghua University(清华大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00234 2025-07-02 cs.LG cs.AI cs.CL 75%

Interpretable AI for Time-Series: Multi-Model Heatmap Fusion with Global Attention and NLP-Generated Explanations

Jiztom Kavalakkatt Francis, Matthew J Darr

机构 * Department of Electrical and Computer Engineering, Iowa State University, Ames, 50011 IA USA(电气与计算机工程系,爱荷华州立大学) Department of Agricultural Biosystem Engineering, Iowa State University, Ames,IA 50011 USA(农业生物系统工程系,爱荷华州立大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02460 2025-06-12 cs.CL cs.AI cs.LG 75%

Code-Switching Curriculum Learning for Multilingual Transfer in LLMs

Haneul Yoo, Cheonbok Park, Sangdoo Yun, Alice Oh, Hwaran Lee

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments To appear in Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15054 2025-06-10 cs.CL cs.AI cs.LG 75%

Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models

Jaturong Kongmanee

机构 * Independent Researcher(独立研究者)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08066 2025-04-14 cs.AI cs.CL cs.LG 75%

The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, David Ha

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10477 2024-02-20 cs.CL cs.AI cs.LG 75%

Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis

Kai Chen, Chunwei Wang, Kuo Yang, Jianhua Han, Lanqing Hong, Fei Mi, Hang Xu, Zhengying Liu, Wenyong Huang, Zhenguo Li, Dit-Yan Yeung, Lifeng Shang, Xin Jiang, Qun Liu

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11656 2026-08-13 cs.LG 新提交 74%

Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models

面向EEG-语言基础模型的语义对齐连续隐式预测建模

Myeong-Ju Cho, Hye-Bin Shin, Seo-Hyun Lee, Seong-Whan Lee

专题命中 其他安全 :alignment(title);分类 cs.LG

AI总结 本文提出BLPM模型,通过CELP编码器与MQSD模块对齐EEG语义,解决EEG基础模型预训练的关键挑战,在多基准任务中实现良好泛化。

Comments 19 pages, 3 figures; supplementary material included

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27739 2026-07-31 cs.IR cs.CL cs.HC 新提交 74%

Measuring Alignment With Reader Highlights Net of Position and Length

通过读者标注的位置和长度网络测量对齐度

Kazuki Nakayashiki, Keisuke Watanabe

专题命中 其他安全 :alignment(title);分类 cs.CL

AI总结 该研究针对上下文压缩评估的混淆问题,提出匹配位置长度的方法,发现语言模型重要性排名对齐度优于人类读者外的基准,且前期某结论无法在当前语料库复现。

Comments 15 pages, 7 tables. Analysis code and de-identified artifacts included as ancillary files; five of six scripts reproduce the paper's numbers from the shipped artifacts alone. Reports claims from our own prior work that this corpus does not reproduce, and lists twelve claims withdrawn during internal adversarial review in Appendix A

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07407 2026-07-28 cs.LG 版本更新 74%

Emergent Symbolic Structure in Health Foundation Models: Extraction, Alignment, and Cross-Modal Transfer

健康基础模型中的涌现符号结构:提取、对齐与跨模态迁移

Gajendra Katuwal, Advait Koparkar, Salar Abbaspourazad, Anshuman Mishra, Sarvesh Kirthivasan

机构 * Apple(苹果公司)

专题命中 其他安全 :alignment(title);分类 cs.LG

AI总结 本文提出一种训练后框架,通过分解冻结嵌入以提取可解释的符号,用于对齐嵌入空间。在PPG和加速度计数据上验证,发现符号能选择性关联健康状况和生理属性,并支持跨模态迁移。

Comments 8 pages, Mechanistic Interpretability Workshop at the 43rd International Conference on Machine Learning, 4 main figures

Journal ref Mechanistic Interpretability Workshop at the 43 rd International Conference on Machine Learning, Seoul, South Korea, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19935 2026-07-23 cs.AI 新提交 74%

MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing

MOF-Sleuth:用于可解释细粒度MOF CIF审核的工具基础奖励对齐

Yu Liu, Zhiwei Yang, Diandian Guo, Kun Peng, Fangfang Yuan, Cong Cao, Chaozhuo Li, Zhiyuan Ma, Yanbing Liu, Guobin Zhao

专题命中 其他安全 :alignment(title);分类 cs.AI

AI总结 研究针对MOF数据库CIF输入错误影响下游结果及人工检查的问题,提出MOF-Sleuth,通过强化引导CIF审核代理的两个模块,利用奖励引导强化学习将工具测量转化为监督,提升检测、归因及解释质量,在多基准测试中性能领先。

详情

展开后加载摘要…

URL PDF HTML 收藏