arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7968 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7968 篇

1307.6569 2015-06-16 astro-ph.HE 78%

Alignment of supermassive black hole binary orbits and spins

M. Coleman Miller, Julian H. Krolik

专题命中 其他安全 :alignment(title,abstract)

Comments 18 pages, 1 figure. Accepted by The Astrophysical Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
1405.2417 2014-05-13 cs.NI 78%

Impact of Two Realistic Mobility Models for Vehicular Safety Applications

Md Habibur Rahman, Mohammad Nasiruddin

专题命中 其他安全 :safety(title,abstract)

Comments 6 pages, 9 figures, ICIEV 2014

Journal ref 3rd International Conference on Informatics, Electronics & Vision (ICIEV 2014)

详情

展开后加载摘要…

URL PDF HTML 收藏
1107.1623 2011-09-26 cond-mat.stat-mech q-bio.OT 78%

Mean-field theory of collective motion due to velocity alignment

Pawel Romanczuk, Lutz Schimansky-Geier

专题命中 其他安全 :alignment(title,abstract)

Comments corrected version, Ecological Complexity (2011) in press

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11221 2024-10-16 cs.LG cs.AI 77%

Multi-objective Reinforcement Learning: A Tool for Pluralistic Alignment

Peter Vamplew, Conor F Hayes, Cameron Foale, Richard Dazeley, Hadassah Harland

专题命中 其他安全 :alignment(title,comments);分类 cs.AI、cs.LG

Comments Accepted for the Pluralistic Alignment workshop at NeurIPS 2024. https://pluralistic-alignment.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06404 2025-06-10 cs.CL cs.AI cs.CY cs.LG 77%

Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights

Sooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao, Xing Xie, JinYeong Bak

机构 * Sungkyunkwan University(釜山大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01405 2025-03-04 cs.LG cs.AI cs.CL cs.CV cs.CY 77%

Representation Engineering: A Top-Down Approach to AI Transparency

Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, Dan Hendrycks

专题命中 其他安全 :safety(abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Code is available at https://github.com/andyzoujm/representation-engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08809 2024-02-08 cs.CL 77%

Interpretability at Scale: Identifying Causal Mechanisms in Alpaca

Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, Noah D. Goodman

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL

Comments NeurIPS 2023 with Author Corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.01478 2022-10-28 cs.CL cs.AI cs.CY cs.LG 77%

When to Make Exceptions: Exploring Language Models as Accounts of Human Moral Judgment

Zhijing Jin, Sydney Levine, Fernando Gonzalez, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, Bernhard Schölkopf

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments NeurIPS 2022 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.02553 2018-01-03 cs.AI cs.SC 77%

Robust Computer Algebra, Theorem Proving, and Oracle AI

Gopal P. Sarma, Nick J. Hay

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

Comments 15 pages, 3 figures

Journal ref Informatica Vol. 41 No. 3 (2017)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09937 2026-08-12 cs.CL cs.CY 新提交 76%

Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory

仔细考量文化:利用文化共识理论分析单文化与多文化场景下的大语言模型对齐

Krishna Pothugunta, John P. Lalor

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.CY

AI总结 本研究利用文化共识理论,分析大语言模型在单/多文化场景下的对齐情况,发现模型存在文化结构误表征问题,该理论可用于区分模型反映人类多样性与算法同质化的情况。

Comments Accepted to ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17629 2026-07-28 cs.LG cs.AI 76%

Learning Chemical Reaction Representation with Reactant-Product Alignment

基于反应物-产物对齐的学习化学反应表示

Kaipeng Zeng, Xianbin Liu, Yu Zhang, Xiaokang Yang, Yaohui Jin, Yanyan Xu

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 本文提出RAlign模型,通过反应物-产物对齐和反应中心感知注意力机制,提升化学反应表示学习的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09936 2026-07-14 cs.LG cs.AI cs.CR 新提交 76%

SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification

SMETA-ZSL:用于零样本威胁分类的语义元对齐

Ivan Alejandro Montoya Sanchez, Anantaa Kotal, Aritran Piplai

机构 * The University of Texas at El Paso(德克萨斯大学艾尔帕索分校)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 研究针对网络安全新威胁无标注数据问题,提出SMETA-ZSL方法,通过对比微调、情景元学习和知识蒸馏等,从重叠语言描述学习语义原型并对齐行为特征,实现跨可见-未见类别的泛化,在7个基准测试中性能远超先前方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01592 2026-07-14 cs.CY cs.CL 交叉投稿 76%

Question Type, Cognitive Load, and CEFR Alignment: Evaluating LLM-Generated EFL Grammar Drill Exercises

问题类型、认知负荷与CEFR对齐:评估LLM生成的EFL语法练习

Steve Woollaston, Brendan Flanagan, Yuko Toyokawa, Hiroaki Ogata

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.CY

AI总结 本研究通过分析日本初中生在语法练习应用中的日志数据,评估了LLM生成的EFL学习内容的教学可行性,揭示了不同问题模态对表现的影响,并验证了CEFR-J语法框架的难度层级。

Comments Under review for the the 34th International Conference on Computers in Education (ICCE 2026). 2jun26: v2 - fixed minor typo

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31394 2026-07-03 cs.LG cs.AI cs.CV q-bio.QM 新提交 76%

Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images

解决AI中的叠加问题以实现可解释性与患者-神经元图像的跨模态对齐

Jisung Park, Seohyeon Kang, Daeun Yoo, Eunsu Lee, Seoin Cho, Wooyeop Choi, Ian Choi, James R. Evan, Daesoo Kim, Sonia Gandhi, Minee L. Choi

机构 * KAIST(韩国科学技术院) Konyang University(建阳大学) Chang Gung University(长庚大学) UCL Queen Square Institute of Neurology & The Francis Crick Institute(伦敦大学学院皇后广场神经病学研究所与弗朗西斯·克里克研究所)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 利用稀疏自编码器解决高维生物数据中神经网络表示空间的叠加问题,恢复几何保真度,并通过Gromov-Wasserstein最优传输实现图像与单细胞RNA测序数据的跨模态对齐。

Comments 10 pages, 7 figures (plus 14 in appendix), 1 table, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04391 2026-07-03 cs.AI cs.CL cs.SI q-bio.NC 版本更新 76%

Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate

心理想象网络显示人类跨群体中心性和聚类对齐,而大型语言模型无法复制

Saurabh Ranjan, Brian Odegaard

机构 * University of Florida(佛罗里达大学)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 本研究通过心理网络分析发现,人类对心理意象的生动性评分在不同文化群体中形成稳定的网络结构,而大型语言模型(LLM)无法复制这种结构,表明人类想象网络根植于具身经验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08016 2026-06-09 cs.CV cs.AI cs.CL 新提交 76%

IEA: Amateur-Friendly Conversational Image Editing Agent via Three Stages of Multitask Alignment

IEA:通过三阶段多任务对齐的业余友好型对话式图像编辑代理

Zichen Zhu, Yuheng Sun, Mingxuan Zhu, Wenjie Ma, Situo Zhang, Zhexiang Wang, Ziyue Yang, Danyang Zhang, Kunyao Lan, Zihan Zhao, Dingye Liu, Siqi Xiang, Lu Chen, Kai Yu

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institution(上海创新研究院) Huawei Technologies Ltd.(华为技术有限公司) Nanyang Technological University(南洋理工大学) Jiangsu Key Lab of Language Computing(江苏省语言计算重点实验室)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 提出IEA对话式图像编辑代理,通过三阶段多任务训练学习操作参数化工具,实现可解释编辑轨迹,在像素距离和ROUGE-L指标上优于基线,用户研究中指令跟随和感知质量表现最佳。

Comments [CVPR 2026 Findings] Our data and code are released at https://github.com/OpenDFM/Image_Edit_Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29458 2026-05-29 cs.CL cs.AI 76%

Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment

面向LLM人格模拟的自适应访谈:基于证据的推理提升决策对齐

Ruoxi Su, Yuhan Liu, Jingyu Hu

机构 * University of Cambridge(剑桥大学) Independent Researcher(独立研究员)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 提出自适应访谈框架,通过结构化三阶段对话收集人格相关信息,并基于访谈记录评估LLM在道德困境场景中模拟个体决策的能力,发现基于后续追问的证据推理能显著提升预测准确性。

Comments 20 pages, 2 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25189 2026-05-26 cs.LG cs.CL 76%

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

方向对齐缓解语言模型强化学习中的奖励黑客问题

Wenlong Deng, Jiaji Huang, Kaan Ozkara, Yushu Li, Christos Thrampoulidis, Xiaoxiao Li, Youngsuk Park

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所) Amazon(亚马逊)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

AI总结 通过分析强化学习更新的几何结构,发现奖励黑客源于优化偏离稳定低维学习轨迹,提出可信方向投影方法约束梯度在干净参考子空间内,延迟捷径利用并保持任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20730 2026-05-21 cs.CL cs.AI 76%

Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning

分布对齐作为设计任务向量在上下文学习中的准则

Jihoon Kwon, Jiwon Choi, Jy-yong Sohn

机构 * Seoul National University(首尔国立大学) Yonsei University(延世大学)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文提出通过分布对齐来设计任务向量,引入了NTP距离作为衡量指标,并开发了线性任务向量方法以提升性能和效率。

Comments 9 pages, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12910 2026-04-23 cs.CL cs.AI 76%

SciCoQA: Quality Assurance for Scientific Paper--Code Alignment

SciCoQA:科学论文与代码对齐的质量保障

Tim Baumgärtner, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(普遍知识处理实验室(UKP实验室)、计算机科学系、德累斯顿技术大学和应用网络安全国家研究中心ATHENE)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文提出SciCoQA数据集,用于评估LLM在科学论文与代码对齐任务中的表现,揭示自动化科学质量保障的关键差距。

Comments Accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04942 2026-04-08 cs.CL cs.AI 76%

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models

TDA-RC:基于任务驱动的知识推理链对齐方法

Jiaquan Zhang, Qigan Sun, Chaoning Zhang, Xudong Wang, Zhenzhen Huang, Yitian Zhou, Pengcheng Zheng, Chi-lok Andy Tai, Sung-Ho Bae, Zeyu Ma, Caiyan Qin, Jinyu Guo, Yang Yang, Hengtao Shen

机构 * School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) College of Professional and Continuing Education, The Hong Kong Polytechnic University(香港理工大学专业及持续教育学院) School of Robotics and Advanced Manufacture, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)机电工程与自动化学院) School of Computing, Kyung Hee University(庆熙大学计算机学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文提出TDA-RC方法,通过拓扑学优化提升大语言模型推理效率与准确性,结合持久同调将不同推理范式统一到拓扑空间中,实现高效且精准的推理链优化。

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19463 2026-04-08 cs.CY cs.AI cs.SI 76%

Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights

对冲与非肯定:量化大语言模型在人权问题上的对齐

Rafiya Javed, Cassandra Parent, Jackie Kay, David Yanni, Abdullah Zaini, Anushe Sheikh, Maribeth Rauh, Walter Gerych, Ramona Comanescu, Iason Gabriel, Marzyeh Ghassemi, Laura Weidinger

机构 * Google Deepmind(谷歌DeepMind) Massachusetts Institute of Technology(麻省理工学院) Independent Researcher(独立研究员) Google(谷歌) AI Accountability Lab, Trinity College Dublin(都柏林圣三一学院人工智能问责实验室)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.CY

AI总结 研究通过系统框架量化LLM在不同群体身份上的对冲与非肯定行为,发现群体身份是主要影响因素,通过引导和正交化技术可有效缓解偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16872 2026-03-19 cs.CL cs.CY 76%

Trust, Safety, and Accuracy: Assessing LLMs for Routine Maternity Advice

信任、安全与准确性:评估LLMs用于常规产科咨询

V Sai Divya, A Bhanusree, Rimjhim, K Venkata Krishna Rao

机构 * National Institute of Technology, Warangal, India(印度战争格尔国家理工学院)

专题命中 其他安全 :safety(title);分类 cs.CL、cs.CY

AI总结 本研究评估了LLMs在提供可靠且易懂的产科信息方面的表现,发现Perplexity在语义上与专家接近,而ChatGPT-4o在文本清晰度和医学术语使用上更优,为偏远地区产科教育提供了可行的AI解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09566 2026-03-12 cs.AI cs.LG 76%

Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search

通过语言模型、性质对齐和战略搜索实现闭环分子发现

Junkai Ji, Zhangfan Yang, Dong Xu, Ruibin Bai, Jianqiang Li, Tingjun Hou, Zexuan Zhu

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 Trio结合语言模型、性质对齐和战略搜索,实现高效且可解释的闭环分子设计,提升药物配体的结合亲和力、药物性和合成可及性。

Comments 30 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06123 2026-01-13 cs.LG cs.AI 76%

Latent Space Communication via K-V Cache Alignment

通过K-V缓存对齐实现潜在空间通信

Lucio M. Dery, Zohar Yahav, Henry Prior, Qixuan Feng, Jiajun Shen, Arthur Szlam

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 本文提出通过学习共享表示空间对齐多模型的k-v缓存,实现模型间高效协作与知识共享,提升整体性能与能力转移。

Comments 15 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24478 2026-01-06 cs.LG cs.AI stat.ME 76%

HOLOGRAPH: Active Causal Discovery via Sheaf-Theoretic Alignment of Large Language Model Priors

HOLOGRAPH:通过sheaf理论对大型语言模型先验进行对齐以实现主动因果发现

Hyunjun Kim

机构 * Korea Advanced Institute of Science \'Ecole Polytechnique F\'ed\'erale de Lausanne (EPFL), Lausanne, Switzerland

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

AI总结 HOLOGRAPH通过sheaf理论对齐大型语言模型先验,实现主动因果发现,提供严谨的数学基础并实现竞争性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00115 2025-12-30 cs.CL cs.AI 76%

Cognitive Alignment in Personality Reasoning: Leveraging Prototype Theory for MBTI Inference

人格推理中的认知对齐:利用原型理论进行MBTI推断

Haoyuan Li, Yuanbo Tong, Yuchen Li, Zirui Wang, Chunhou Liu, Jiamou Liu

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 ProtoMBTI通过原型理论与认知对齐,提升文本人格推断的准确性和泛化能力。

Comments The authors have decided to withdraw this version to substantially revise and extend the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15545 2025-12-30 cs.CL cs.AI 76%

TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs

TokenTiming: 一种适用于通用推测解码模型对的动态对齐方法

Sibo Xiao, Jinyuan Fu, Zhongle Xie, Lidan Shou

机构 * Zhejiang University(浙江大学)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 TokenTiming通过动态对齐方法实现通用推测解码,无需重新训练即可处理不同词汇模型,提升LLM推理效率1.57倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04834 2025-11-13 cs.LG cs.AI cs.CV 76%

Prompt-Based Safety Guidance Is Ineffective for Unlearned Text-to-Image Diffusion Models

Jiwoo Shin, Byeonghu Na, Mina Kang, Wonhyeok Choi, Il-Chul Moon

机构 * KAIST(韩国科学技术院)

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2025 Workshop on Generative and Protective AI for Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14158 2025-11-11 cs.CL cs.LG 76%

Temporal Alignment of Time Sensitive Facts with Activation Engineering

Sanjay Govindan, Maurice Pagnucco, Yang Song

机构 * University of New South Wales(新南威尔士大学)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏