arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-05-05 至 2026-05-05 共收录 54 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 22 篇

2605.00839 2026-05-05 cs.AI cs.LG 62%

2026 Roadmap on Artificial Intelligence and Machine Learning for Smart Manufacturing

2026 年人工智能与机器学习在智能制造中的路线图

Jay Lee, Hanqi Su, Marco Macchi, Adalberto Polenghi, Wei Wu, Zhiheng Zhao, George Q. Huang, Kiva Allgood, Devendra Jain, Benedikt Gieger, Vibhor Pandhare, Soumyabrata Bhattacharjee, Ram Mohril, Lingbao Kong, Qiyuan Wang, Xinlan Tang, Sungjong Kim, Chan Hee Park, Byeng D. Youn, Guo Dong Goh, Xi Huang, Wai Yee Yeong, Yung C Shin, He Zhang, Zitong Wang, Fei Tao, Jagjit Singh Srai, Satyandra K. Gupta, Byung Gun Joung, Albin John, John W. Sutherland, Sang Won Lee, Olga Fink, Vinay Sharma, Faez Ahmed, Wei Chen, Mark Fuge, Arild Waaler, Martin G. Skjæveland, Dimitris Kyritsis, Wei Chen, VispiNevile Karkaria, Yi-Ping Chen, Ying-Kuan Tsai, Joseph Cohen, Xun Huan, Jing Lin, Liangwei Zhang, Gregory W. Vogl, Aaron W. Cornelius, Xiaodong Jia, Dai-Yan Ji, Takanobu Minami, Ruoxin Wang

机构 * Center for Industrial Artificial Intelligence, Department of Mechanical Engineering, University of Maryland, College Park(工业人工智能中心,机械工程系,马里兰大学College Park分校) Department of Management, Economics and Industrial Engineering, Politecnico di Milano(管理、经济与工业工程系,米兰理工学院) Department of Industrial and Systems Engineering, The Hong Kong Polytechnic University(工业与系统工程系,香港理工大学) Centre for Advanced Manufacturing & Supply Chains, World Economic Forum(先进制造与供应链研究中心,世界经济论坛) Department of Mechanical Engineering, Indian Institute of Technology Indore(机械工程系,印度理工学院Indore分校) Future Information Innovative College, Fudan University(未来信息创新学院,复旦大学) Department of Mechanical Engineering, Seoul National University(机械工程系,首尔国立大学) Department of Mechanical and Information Engineering, University of Seoul(机械与信息工程系,首尔大学) Onepredict Corp.(Onepredict公司) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学) Singapore Centre for 3D Printing, Nanyang Technological University(新加坡3D打印中心,南洋理工大学) Mechanical Engineering, Purdue University(机械工程系,普渡大学) Digital Twin International Research Center, International Institute for Interdisciplinary and Frontiers, Beihang University(数字孪生国际研究中心, interdisciplinary and Frontiers 国际研究院,北京航空航天大学) School of Automation Science and Electrical Engineering, Beihang University(自动化科学与电气工程学院,北京航空航天大学) Department of Engineering, University of Cambridge(工程系,剑桥大学) Center for Advanced Manufacturing, University of Southern California(先进制造中心,南加州大学) School of Sustainability Engineering and Environmental Engineering, Purdue University(可持续工程与环境工程系,普渡大学) School of Mechanical Engineering, Sungkyunkwan University(机械工程系,全南大学) Intelligent Maintenance and Operations Systems, EPFL(智能维护与运营系统,苏黎世联邦理工学院) Department of Mechanical Engineering, Massachusetts Institute of Technology(机械工程系,麻省理工学院) J. Mike Walker ’66 Department of Mechanical Engineering, Texas A&M University(J. Mike Walker ’66 机械工程系,德克萨斯A&M大学) Department of Mechanical and Process Engineering, ETH Zürich(机械与工艺工程系,苏黎世联邦理工学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨人工智能与机器学习在智能制造中的发展现状与未来方向,涵盖基础理论、应用领域及新兴技术,旨在推动创新与产业应用。

Comments This paper has been accepted for publication in the Journal Machine Learning: Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05534 2026-05-05 cs.LG cs.AI 62%

A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima

稀疏字典学习在机制可解释性中的统一理论:分段双凸性与虚假极值

Yiming Tang, Harshvardhan Saini, Zhaoqian Yao, Zheng Lin, Yizhen Liao, Jingyi Cui, Yisen Wang, Mengnan Du, Dianbo Liu

机构 * National University of Singapore(新加坡国立大学) Indian Institute of Technology, Dhanbad(印度德里理工学院) Chinese University of Hong Kong(香港中文大学) Hong Kong University of Science and Technology(香港科学与技术大学) Peking University(北京大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出稀疏字典学习的统一理论,揭示其分段双凸性及虚假极值问题,通过线性表示基准揭示路径学,提出特征锚定技术提升特征恢复性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01611 2026-05-05 cs.CY cs.AI cs.LG 60%

The Case for ESM3 as a General-Purpose AI Model with Systemic Risk Under the EU AI Act

ESM3作为具有系统性风险的通用人工智能模型在欧盟人工智能法案中的案例

Taro Qureshi, Jacob Griffith, Koen Holtman, Marcel Mir Teijeiro, Ze Shen Chin, Rokas Gipiškis

机构 * AI Standards Lab(人工智能标准实验室) Vilnius University(维尔纽斯大学) Northeastern University London(伦敦东北大学) London School of Economics(伦敦政治经济学院)

专题命中 安全评测 :分类 cs.AI、cs.CY、cs.LG;safety(comments);AI safety(comments)

AI总结 本文探讨ESM3等前沿生物基础模型在欧盟人工智能法案下的监管问题,分析其是否受通用人工智能模型系统性风险义务约束,并提出改进建议。

Comments 8 pages, 1 figure, Technical AI Safety Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01546 2026-05-05 cs.NI cs.AI 57%

6G Needs Agents: Toward Agentic AI-Native Networks for Autonomous Intelligence

6G需要智能体:迈向面向自主智能的智能体AI原生网络

Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah

机构 * Department of Computer and Network Engineering, United Arab Emirates University, UAE(计算机与网络工程系,阿联酋大学) Research Institute for Digital Future, Khalifa University, UAE(数字未来研究院,哈利法大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出基于大语言模型的智能体架构,用于构建AI原生6G网络,通过分层推理和分布式多智能体系统平衡性能与效率,揭示了模型异构部署和量化影响的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01483 2026-05-05 cs.CV cs.AI 57%

Research on Vision-Language Question Answering Models for Industrial Robots

工业机器人中视觉-语言问答模型的研究

Ping Li, Bartlomiej Brzozka

机构 * Shandong Management University(山东管理大学) Maria Curie-Sklodowska University(玛丽·居里-斯洛多夫斯卡大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出一种层次交叉模态融合模型,解决工业机器人中语义模糊、复杂环境布局和领域术语的问题,通过多级特征融合提升问答可靠性。

Comments 8 Pages, 5 figures

Journal ref Brzozka,B.Machine Learning Algorithms in Predicting College Students' Grades:A Review.Journal of Applied Automation Technologies,2025,3,1-12

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01302 2026-05-05 cs.CL cs.IR 57%

Beyond Semantic Relevance: Counterfactual Risk Minimization for Robust Retrieval-Augmented Generation

超越语义相关性:为鲁棒检索增强生成的反事实风险最小化

Peiyang Liu, Qiang Yan, Ziqiang Cui, Di Liang, Xi Wang, Wei Ye

机构 * National Engineering Research Center for Software Engineering, Peking University(软件工程国家工程研究中心,北京大学) City University of Hong Kong(香港城市大学) Tencent Technology(腾讯科技) Peking University(北京大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

AI总结 本文提出CoRM-RAG框架,通过因果干预模拟用户偏差来提升检索鲁棒性,在对抗性场景中优于传统检索器和重排序器。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00957 2026-05-05 cs.IR cs.AI 57%

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation

我不知道--迈向具有确定性意识的适当信任

Daan Di Scala, Maaike de Boer, Pınar Yolum

机构 * TNO Netherlands Organisation for Applied Scientific Research, Department Data Science(荷兰应用科学研究院,数据科学部门) Utrecht University, Department of Information and Computing Sciences(乌得勒支大学,信息与计算科学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出CERTA系统,通过结合问题、上下文和答案的相关性来反映不确定性,以建立适当的信任。研究创建了Certainty Benchmark,并通过实验验证了CERTA在减少过度同意和提供谨慎行为方面的有效性。

Comments To be published in VALE 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27033 2026-05-05 cs.LG eess.SP 57%

Cross-Subject Generalization for EEG Decoding: A Survey of Deep Learning Methods

跨受体解码的通用性:深度学习方法的综述

Taida Li, Yujun Yan, Fei Dou, Wenzhan Song, Xiang Zhang

机构 * Department of Computer Science, University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校计算机科学系) Department of Computer Science, Dartmouth College(达特茅斯学院计算机科学系) School of Computing, University of Georgia(佐治亚大学计算科学学院) School of Electrical and Computer Engineering, University of Georgia(佐治亚大学电气与计算机工程学院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

AI总结 本文综述了深度学习方法在跨受体解码中的应用,分析了多源领域问题的评估协议,并系统分类了特征对齐、对抗学习等方法,探讨了理论限制和EEG基础模型的发展。

Comments Accepted manuscript in Progress in Biomedical Engineering. Minor update: corrected author affiliation in comment

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24814 2026-05-05 cs.CV cs.AI 57%

Deep Feature Optimization for Enhanced Fish Freshness Assessment

深度特征优化用于增强鱼类新鲜度评估

Phi-Hung Hoang, Nam-Thuan Trinh, Van-Manh Tran, Thi-Thu-Hong Phan

机构 * Department of Artificial Intelligence, FPT University(人工智能系,FPT大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 本文提出一种三阶段框架,通过优化深度特征提升鱼类新鲜度评估的准确性与可靠性,实验表明其在鱼类眼睛新鲜度数据集上达到85.99%的准确率。

Comments 39 pages; 10 tables; 9 figures

Journal ref Ecological Informatics, 95, 103711, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05284 2026-05-05 cs.LG 57%

Analyzing Adversarial Inputs in Deep Reinforcement Learning

分析深度强化学习中的对抗输入

Davide Corsi, Guy Amir, Guy Katz, Alessandro Farinelli

机构 * University of California, Irvine, USA(加州大学尔布雷特分校) Cornell University, USA(康奈尔大学) The Hebrew University of Jerusalem, Israel(特拉维夫大学) University of Verona, Italy(威尼斯大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

AI总结 本文通过形式验证方法分析深度强化学习中对抗输入的特性,提出Adversarial Rate指标,用于评估对抗输入对DRL策略的影响,并提供工具和算法以缓解训练DRL网络的脆弱性。

Comments Accepted to AISoLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00927 2026-05-05 q-bio.OT 50%

BioVeil MATRIX: Uncovering and categorizing vulnerabilities of agentic biological AI scientists

BioVeil MATRIX: 洞察和分类代理生物AI科学家中的漏洞

Kimon Antonios Provatas, Avery Self, Ioannis Mouratidis, Ilias Georgakopoulos-Soares

专题命中 安全评测 :safety(abstract)

AI总结 研究探讨了代理生物AI科学家在多学科科学流程中的应用,指出现有安全评估无法覆盖双用途应用风险,提出BioVeil MATRIX作为防御分类体系,用于系统化分类生物安全风险。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 3 篇

2605.01147 2026-05-05 cs.AI 89%

Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

位置:代理AI的安全性与公平性取决于交互拓扑,而非模型规模或对齐

Tanav Singh Bajaj, Nikhil Singh, Karan Anand, Eishkaran Singh

机构 * Department of Computer Science, University of British Columbia, Vancouver, Canada(英属哥伦比亚大学计算机科学系) Department of Artificial Intelligence, IIT Hyderabad, Hyderabad, India(印度海得拉巴理工学院人工智能系) Amazon, Delhi, India(印度德里亚马逊公司)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(title,abstract);AI safety(abstract);分类 cs.AI

AI总结 本文指出代理AI的安全性由交互拓扑决定,而非模型规模或对齐程度。研究揭示了顺序不稳定、信息级联和功能崩溃等拓扑驱动的问题,并强调需通过动态系统视角评估安全性和公平性。

Comments 18 pages, 8 figures. Position paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01229 2026-05-05 cs.LG cs.CL 62%

Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation

大规模多语言神经机器翻译中的注意力 sinks:发现、分析与缓解

Hillary Mutisya, John Mugane

机构 * Harvard University(哈佛大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 研究发现神经机器翻译中跨注意力模式存在注意力 sinks,非内容token占据大部分注意力质量,影响相似性评估,提出内容过滤方法以恢复语言信号。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01451 2026-05-05 cs.CL 61%

Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models

对基于AI的紧急警务调度中的种族偏见进行审计:对十一种大型语言模型的跨语言评估

William Guey, Wei Zhang, Pierrick Bougault, Yi Wang, Bertan Ucar, Vitor D. de Moura, José O. Gomes

机构 * Department of Industrial Engineering, Tsinghua University(清华大学工业工程系) School of Social Sciences, Tsinghua University(清华大学社会科学部) Department of Industrial Engineering, Federal University of Rio de Janeiro(里约热内卢联邦大学工业工程系)

专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.CL

AI总结 本文通过跨语言框架评估11种模型,在19800个输出中发现当事件严重性模糊时种族偏见系统性出现,但当操作优先级由通话内容确定时偏见消失。偏见程度因种族轴而异,宗教外观影响最大,性别次之,种族最小。语言间偏见转移不一致,性别偏见在中文中放大,种族偏见在英文中更明显。

Comments 26 pages, 7 figures. Submitted to Humanities and Social Sciences Communications (Nature) collection on Artificial Intelligence and Emerging Technologies in Public Safety. Code and data: https://github.com/williamguey/llmdispatchbias

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 10 篇

2605.00842 2026-05-05 cs.AI cs.LG 73%

Understanding Emergent Misalignment via Feature Superposition Geometry

通过特征叠加几何理解涌现对齐问题

Gouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo

机构 * The University of Tokyo(东京大学) Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 研究通过特征叠加几何解释了微调窄任务导致有害行为的机制,发现有害特征在几何上更接近,并通过过滤接近毒特征的数据减少对齐问题。

Comments Accepted to ACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01153 2026-05-05 cs.HC 67%

Toward a Unified Framework for Collaborative Design of Human-AI Interaction

迈向人机交互协作设计的统一框架

Ankur Bhatt, Sven Mayer

专题命中 其他安全 :alignment(abstract);safety(abstract)

AI总结 本文提出统一框架整合多模态对齐、交互导向的可解释性及用户代理,通过协作设计与增强现实仓库机器人场景验证,提升人机交互透明度与用户控制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01609 2026-05-05 cs.LG cs.AI 62%

Concepts Whisper While Syntax Shouts: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations

概念低语而语法高呼:频谱反集中与Transformer表示的双几何

Pratyush Acharya, Nuraj Rimal, Habish Dhakal

机构 * Pratyush Acharya Nuraj Rimal Habish Dhakal

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 研究通过跨语言概念传输测试,发现因果内积在频谱正则化下无显著差异,揭示了残差流中反集中现象及静态解嵌行对比的高方差集中特性,支持语法优先编码于高方差子空间。

Comments 25 pages, 16 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01167 2026-05-05 cs.LG cs.AI 62%

Minimizing Collateral Damage in Activation Steering

最小化激活引导中的附带损害

Tam Nguyen, Tu Anh Nguyen, Sina Alemohammad, Richard G. Baraniuk

机构 * Department of Electrical \& Computer Engineering, Rice University, Houston, USA Department of Computational Applied Mathematics, Rice University, Houston, USA Department of Electrical \& Computer Engineering, The University of Texas at Austin, Austin, USA

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于约束优化的框架,通过数学形式化附带损害并优化激活变化,以更精确地控制大语言模型行为,同时减少对无关任务性能的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17815 2026-05-05 cs.HC cs.CL cs.CY 62%

Navigating the Conceptual Multiverse

在概念多元宇宙中导航

Andre Ye, Jenny Y. Huang, Alicia Guo, Rose Novick, Tamara Broderick, Mitchell L. Gordon

机构 * MIT EECS(麻省理工学院电子工程与计算机科学系) UW Allen School of CSE(华盛顿大学Allen学院计算机科学与工程系) UW Department of Philosophy(华盛顿大学哲学系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 本文提出概念多元宇宙,通过交互系统让用户透明地检查和改变概念决策,提升问题理解。在三个领域中,该系统帮助参与者更清晰地把握问题,如哲学学生改写论文、对齐标注者深入分析用户意图、诗人发现创作模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01325 2026-05-05 cs.CV cs.LG 57%

Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance

通过格罗莫夫-瓦瑟斯坦距离重新思考VLM中的模型选择

Muyang Li, Yucheng Liu, Jianbo Ma, Elliot Osborne, Bo Han, Tongliang Liu

机构 * Sydney AI Centre, The University of Sydney(悉尼人工智能中心,悉尼大学) Dolby Laboratories(杜比实验室) TMLR Group, Hong Kong Baptist University(香港 Baptist 大学 TMLR 团体)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文通过系统实验探讨视觉编码器选择的关键因素,发现模态间结构相似性通过格罗莫夫-瓦瑟斯坦距离衡量能提升VLM性能预测。

Comments Accepted as Highlight publication for CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01245 2026-05-05 cs.HC cs.AI 57%

The Garden of Forking Paths: Narrative Arc-Conditioned Gameplay Planning

分叉花园:叙事弧条件下的游戏玩法规划

Yunge Wen, Chenliang Huang, Hangyu Zhou, Zhuo Zeng, Chun Ming Louis Po, Julian Togelius, Timothy Merino, Sam Earle

机构 * New York University(纽约大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出Forking Garden框架,通过用户提供的故事情节生成分支游戏,采用弧引导约束算法构建 dungeon 图,实现游戏元素的多模态对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07885 2026-05-05 cs.CR cs.AI cs.SE 57%

False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models

壳中的假朋友:揭示大型语言模型中的表情符号语义混淆

Weipeng Jiang, Xiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao Shen, Yang Liu

机构 * Xi’an Jiaotong University(西安交通大学) Nanyang Technological University(南洋理工大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 研究揭示大型语言模型对表情符号的语义混淆问题,通过构建数据集发现平均混淆率超38%,且多数混淆响应导致安全风险,呼吁开发有效缓解方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01055 2026-05-05 cs.SE cs.AI cs.CR 57%

Towards Agentic Runtime Healing

面向代理的运行时自愈

Zhensu Sun, Haotian Zhu, Bowen Xu, Xiaoning Du, Li Li, David Lo

机构 * Singapore Management University(新加坡国立大学) North Carolina State University(北卡罗来纳州立大学) Monash University(墨尔本大学) Beihang University(北航)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 本文提出利用大语言模型实现动态生成错误处理策略,通过Healer框架在四个代码数据集上验证了其在运行时错误恢复中的有效性,展示了LLM在自愈系统中的潜力。

Comments Accepted by CACM

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01190 2026-05-05 astro-ph.IM astro-ph.EP astro-ph.SR 50%

Statistically Significant Linear Alignments Among High-Confidence Transient Candidates on POSS-I Photographic Plates

在POSS-I摄影底板上高置信度暂现候选体中统计显著的线性对齐

Brian Doherty

专题命中 其他安全 :alignment(abstract)

AI总结 研究发现POSS-I摄影底板上高置信度暂现候选体中存在统计显著的线性对齐和异常空间聚类,通过机器学习分类器筛选出107,875个候选体,发现7块底板上的对齐源超过蒙特卡洛期望,排除了持续发光天体的可能性。

Comments 20 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏