Soft Alignment Objectives for Robust Adaptation of Language Generation
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI
Comments Annual Meeting of The ACL 2023: Main conference long paper
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI
Comments Annual Meeting of The ACL 2023: Main conference long paper
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG
Comments Accepted to ICML 2023 and CVPR4XAI workshop 2023
专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG
Comments 40 pages
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.LG
Comments Fixed typos
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG
Comments ICPR 2022
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG
专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG
Comments International Science and Innovation Congress 2019, pp. 643-655, 13 pages, 10 figures
专题命中 其他安全 :alignment(title);分类 cs.AI、cs.CY
量化理论上的AI对齐保证:贝叶斯说服中的接收者效用界
机构 * Cornell University(康奈尔大学)
专题命中 其他安全 :alignment(title,comments);分类 cs.AI
AI总结 通过贝叶斯说服模型,研究AI发送者优化错位目标时,人类接收者仍能获得多少有用信息,证明接收者效用比不超过3/2,并给出紧性下界。
Comments 12 pages, EC 2026 Poster and EC 2026 Incentive-Based AI Alignment Workshop Poster
机构 * Department of Linguistics and Modern Languages, The Chinese University of Hong Kong(语言学与现代语言系,香港中文大学) ; Brain and Mind Institute, The Chinese University of Hong Kong(脑与心智研究所,香港中文大学)
专题命中 其他安全 :alignment(title,comments);分类 cs.CL
Comments Hanlin Wu, Xufeng Duan, and Zhenguang Cai. 2025. Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 135-143, Albuquerque, New Mexico, USA. Association for Computational Linguistics. https://aclanthology.org/2025.cmcl-1.18/
Journal ref In Proceedings of CMCL, pages 135-143, ACL (2025)
并非所有大语言模型的推理都能在思维链中体现
机构 * New York University(纽约大学) ; University of Maryland(马里兰大学) ; TogetherAI
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究探讨大语言模型输出令牌是否体现所有推理,发现前沿模型存在利用无关填充令牌提升合成推理任务性能的不可见推理现象,评估多个模型,揭示填充令牌益处因模型和令牌而异,还表明其能服务隐藏目标,且强化学习等方法无法使填充令牌益处在测试时持续。
CuMA: 通过人口统计感知的适配器混合使大语言模型与稀疏文化价值观对齐
机构 * Southeast University(东南大学) ; ByteDance Inc.(字节跳动公司) ; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用重点实验室(东南大学),中华人民共和国教育部,中国)
专题命中 其他安全 :alignment(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG
AI总结 提出CuMA框架,通过人口统计感知路由将冲突梯度分离到专家子空间,解决密集模型在多文化对齐中的均值崩溃问题,在WorldValuesBench等基准上取得最优性能。
Comments ACL 2026 Main
检查你的大语言模型的秘密词典!五行代码揭示你的大语言模型学到了什么(包括它不应该学到的)
机构 * Mgnite Inc.(Mgnite公司)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 通过对lm_head权重矩阵进行奇异值分解(仅需五行PyTorch代码且无需模型推理),直接从模型权重中揭示可解释的语义子空间,并发现模型训练数据组成和策展哲学。
对Llama3-8b-Instruct自生成文本识别能力的检查与控制
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本研究探讨了LLM是否能识别自身生成的文本,发现Llama3-8b-Instruct模型能够区分自身输出与人类输出,并通过残差流中的特定向量控制其行为和感知,揭示了模型自我归属的认知机制。
Comments 10 pages, 13 figs, 2 tables, accepted as conference paper to ICLR 2025
Journal ref The Thirteenth International Conference on Learning Representations (ICLR 2025)
计划是什么?LLMs中隐式规划的度量及其在押韵生成和问答中的应用
机构 * HPI / University of Potsdam(HPI/波茨坦大学) ; Utrecht University(乌特勒支大学) ; Google DeepMind(谷歌DeepMind)
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出简单方法评估LLM隐式规划,通过押韵生成和问答案例展示其可扩展性,发现隐式规划在1B参数模型中普遍存在,为AI安全提供新视角。
Comments 41 pages, 34 figures, Accepted at ICLR 2026, Code available at https://github.com/Jim-Maar/implicit-planning-in-llms
刻画模型内禀技能
机构 * Virginia Tech(弗吉尼亚理工学院)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出模型内禀技能的刻画方法,通过从序列激活中恢复紧凑正交基,实现行为变化轴的自组织,验证了在推理和安全对齐中的有效性,优于人类定义的技能。
Comments We argue that when the goal is to intervene on model behavior, skill characterization should be *model-native*: grounded in the model's own representations rather than imposed through external ontologies
DiaBlo: 对角块足以用于微调
机构 * University at Albany, SUNY(纽约州立大学阿尔巴尼分校) ; IBM T. J. Watson Research Center(IBM 汤普逊·杰·沃森研究中心)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 DiaBlo是一种仅更新模型权重矩阵对角块的参数高效微调方法,通过消除低秩矩阵乘积需求,实现稳定收敛和高效训练。
Comments Accepted by ICLR 2026
迭代部署提升大语言模型的规划能力
机构 * University of Oxford(牛津大学) ; AI Sequrity Company(AI安全公司) ; UFRGS(乌拉圭联邦大学)
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 通过迭代部署大语言模型,利用用户编纂的数据提升规划能力,展现隐含奖励函数的强化学习机制,具有AI安全和训练制度替代的双重意义。
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 46 pages, 17 figures, 26 tables. Submitted for publication. for associated blog post, see https://pradyut3501.github.io/lora-spur-corr/
机构 * University of Maryland(马里兰大学) ; Tsinghua University(清华大学)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments COLM 2025
机构 * Department of Electrical and Computer Engineering, Iowa State University, Ames, 50011 IA USA(电气与计算机工程系,爱荷华州立大学) ; Department of Agricultural Biosystem Engineering, Iowa State University, Ames,IA 50011 USA(农业生物系统工程系,爱荷华州立大学)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 13 pages
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments To appear in Findings of ACL 2025
机构 * Independent Researcher(独立研究者)
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by ICLR 2024
面向EEG-语言基础模型的语义对齐连续隐式预测建模
专题命中 其他安全 :alignment(title);分类 cs.LG
AI总结 本文提出BLPM模型,通过CELP编码器与MQSD模块对齐EEG语义,解决EEG基础模型预训练的关键挑战,在多基准任务中实现良好泛化。
Comments 19 pages, 3 figures; supplementary material included
通过读者标注的位置和长度网络测量对齐度
专题命中 其他安全 :alignment(title);分类 cs.CL
AI总结 该研究针对上下文压缩评估的混淆问题,提出匹配位置长度的方法,发现语言模型重要性排名对齐度优于人类读者外的基准,且前期某结论无法在当前语料库复现。
Comments 15 pages, 7 tables. Analysis code and de-identified artifacts included as ancillary files; five of six scripts reproduce the paper's numbers from the shipped artifacts alone. Reports claims from our own prior work that this corpus does not reproduce, and lists twelve claims withdrawn during internal adversarial review in Appendix A
健康基础模型中的涌现符号结构:提取、对齐与跨模态迁移
机构 * Apple(苹果公司)
专题命中 其他安全 :alignment(title);分类 cs.LG
AI总结 本文提出一种训练后框架,通过分解冻结嵌入以提取可解释的符号,用于对齐嵌入空间。在PPG和加速度计数据上验证,发现符号能选择性关联健康状况和生理属性,并支持跨模态迁移。
Comments 8 pages, Mechanistic Interpretability Workshop at the 43rd International Conference on Machine Learning, 4 main figures
Journal ref Mechanistic Interpretability Workshop at the 43 rd International Conference on Machine Learning, Seoul, South Korea, 2026
MOF-Sleuth:用于可解释细粒度MOF CIF审核的工具基础奖励对齐
专题命中 其他安全 :alignment(title);分类 cs.AI
AI总结 研究针对MOF数据库CIF输入错误影响下游结果及人工检查的问题,提出MOF-Sleuth,通过强化引导CIF审核代理的两个模块,利用奖励引导强化学习将工具测量转化为监督,提升检测、归因及解释质量,在多基准测试中性能领先。