Personal Universes: A Solution to the Multi-Agent Value Alignment Problem
专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI
专题命中 其他安全 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.LG
Comments AAAI Alignment Track 2025 Poster
语言模型对齐的理论极限
机构 * Apple(苹果公司)
专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.CY、cs.LG
AI总结 研究语言模型对齐的理论极限,通过推导KL散度预算下的最大预期奖励增益,揭示了KL正则化对齐的信息论限制,并证明了奖励融合能缓解奖励黑客问题。
为何智能体在压力下妥协安全
机构 * Guangdong Provincial Key Laboratory of Brain-inspired Intelligent Computation, Department of Computer Science and Engineering, Southern University of Science and Technology(广东省脑启发智能计算重点实验室,计算机科学与工程系,南方科技大学)
专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.CY
AI总结 本文研究了智能体在复杂环境中为追求目标而妥协安全的问题,揭示了'智能体压力'概念,并提出通过压力隔离等策略缓解这一现象。
Comments Accepted by ACL 2026 Findings; 18 pages, 5 figures
机构 * Nanyang Technological University(南洋理工大学) ; National University of Singapore(新加坡国立大学) ; The Hong Kong Polytechnic University(香港理工大学) ; A*STAR(科技研究局) ; Southern University of Science and Technology(南方科技大学) ; University of Science and Technology of China(中国科学技术大学) ; The Pennsylvania State University(宾夕法尼亚州立大学) ; TeleAI ; Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Zhejiang University(浙江大学) ; Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; Renmin University of China(中国人民大学) ; University of California, San Diego(加州大学圣地亚哥分校) ; Tencent(腾讯) ; Georgia Institute of Technology(佐治亚理工学院) ; Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments The first two authors contribute equally to this work and are listed in alphabetical order
专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 其他安全 :alignment(title,abstract);harmlessness(abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Short version accepted as a Tiny Paper at the International Conference on Learning Representations (ICLR) 2024. Long version accepted to the Conference on Empirical Methods in Natural Language Processing (EMNLP) 2024 Findings
机构 * Aran Nayebi(独立研究者)
专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG
Comments 21 pages, 1 figure, 1 table. To appear in AAAI 2026 Special Track on AI Alignment (oral)
专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.CY
Comments Pluralistic Alignment Workshop at NeurIPS 2024
专题命中 其他安全 :safety(title);AI safety(title);分类 cs.AI
在执行轨迹上进行推理时间对齐的工具
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG
AI总结 本文研究了在执行轨迹上进行推理时间对齐的工具设计,通过任务分解和引导执行机制来提高长期性能,发现工具设计中分解和引导的复杂性并不总是带来更好的结果,提出了任务分解和引导执行的两种机制,并通过合成实验和实际终端代理基准验证了这些发现。
SafeAnchor:在大语言模型连续领域适应中防止累积安全性侵蚀
机构 * The University of Hong Kong(香港大学) ; Brain Investing Limited
专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.LG
AI总结 SafeAnchor通过识别低秩安全子空间并约束梯度更新,有效防止连续领域适应中的安全性侵蚀,实验显示其在多个领域任务中保持了较高的安全性对齐度。
Comments 16 pages (12 main + 4 appendix), 2 figures, 12 tables
我难以相信它不稳健:在嵌入漂移下安全分类器的灾难性崩溃
机构 * Independent(独立研究者) ; Meta AI ; AWS Generative AI Innovation Center, Amazon Web Services(AWS生成式AI创新中心,亚马逊网络服务) ; Northeastern University, Seattle, WA, USA(东北大学,西雅图,华盛顿州,美国) ; Stanford University(斯坦福大学)
专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.CL、cs.LG
AI总结 研究发现嵌入漂移导致安全分类器性能大幅下降,揭示了生产AI安全架构的脆弱性并挑战了安全机制的转移假设。
Comments Accepted at the ICBINB: Where LLMs Need to Improve workshop at ICLR 2026. 12 pages and 3 Figures
专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI
Comments 20 pages, 7 figures, 4 tables
专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI、cs.CY
Comments 14 pages and 2 figures + 6 pages for references and appendices
专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI、cs.CY
用于持续指令微调的渐进式多模态对齐
机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences(中国科学院脑科学与智能技术卓越创新中心) ; Kyoto University(京都大学) ; Migu Culture Technology Co.,Ltd.(咪咕文化科技有限公司) ; State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与类脑智能技术国家重点实验室)
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI
AI总结 针对多模态持续指令微调中投影器级遗忘问题,提出渐进式多模态对齐框架PMA,以亚线性参数增长平衡稳定性与可塑性,在多基准实验中提升了现有方法性能且适配多种MLLM主干。
Comments Accepted by ACM MM2026
抱歉,司机,恐怕我不能这么做:评估LLMs在汽车环境中的安全性
机构 * UKRI AI Centre for Doctoral Training in Safe Artificial Intelligence Systems (SAINTS)(英国研究理事会安全人工智能系统博士培训中心(SAINTS)) ; University of York(约克大学) ; Jaguar Land Rover(捷克·陆罗恩)
专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI
AI总结 本文从安全保证角度评估了将LLMs集成到汽车控制任务中的现有框架,指出其面临概念和具体挑战,并通过案例研究提出未来保障机制。
Comments Accepted at the Dependable AI in Embedded Systems (DAIES) Workshop at SAFECOMP 2026; 15 pages, 3 figures, 2 tables
通过智能体对话危害识别分析增强操作安全性
机构 * Oak Ridge National Laboratory(橡树岭国家实验室)
专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI
AI总结 提出HAZDIAL框架,利用结构化多智能体多轮对话(对抗性辩论与建设性讨论)改进基于NLP的危害识别质量,并通过算法优化智能体交互,实验证明优于单次基线方法。
异质双曲流形上的树间模态对齐
机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学) ; Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University(广东机器感知与智能计算实验室,深圳MSU-BIT大学) ; Department of Electrical and Computer System Engineering, Monash University(电子与计算机系统工程系,墨尔本大学)
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
AI总结 提出一种在异质双曲流形上对齐图像和文本树状层次特征的方法,通过交叉注意力提取视觉层次特征、异质流形嵌入及KL距离度量学习中间流形,在开放集分类任务中优于基线。
Comments Published as a conference paper at ICLR 2026
Journal ref The Fourteenth International Conference on Learning Representations (ICLR 2026), Rio de Janeiro, Brazil, 2026
模型规范中期训练:改进对齐训练的泛化能力
机构 * Anthropic
专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI
AI总结 提出模型规范中期训练(MSM),通过在预训练后、对齐微调前用合成文档训练模型学习规范内容,从而引导模型从后续演示数据中泛化,有效降低代理失调率并提升对齐泛化能力。
通过协作对齐的联邦LoRA微调大型语言模型
机构 * School of Computing & Data Science, The University of Hong Kong(计算与数据科学学院,香港大学)
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
AI总结 本文研究了在联邦学习环境下使用LoRA进行参数高效微调的问题,提出了一种名为CLAIR的框架,通过结构低秩加块稀疏分解来恢复共享LoRA子空间并检测污染客户端,从而在噪声情况下实现精确恢复,并在不同条件下实现稳定和一致的协作集恢复。
ProjLens: 揭示项目器在多模态模型安全中的作用
机构 * University of Science and Technology of China(中国科学技术大学) ; Beijing University of Aeronautics and Astronautics(北京航空航天大学) ; Nanyang Technological University(南洋理工大学) ; Southwest Jiaotong University(西南交通大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI
AI总结 ProjLens通过分析多模态大语言模型中的后门攻击机制,揭示了项目器在安全漏洞中的关键作用,发现后门注入参数编码于低秩子空间,并通过实验验证了激活机制的差异。
Comments 18 pages ,15 figures