arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-05-25 至 2026-05-25 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 4 篇

2605.22963 2026-05-25 cs.CL cs.AI 81%

Graph Alignment Topology as an Inductive Bias for Grounding Detection

图对齐拓扑作为接地检测的归纳偏置

Paul Landes, Pranav Herur, Adam Cross, Jimeng Sun

机构 * Department of Pediatrics, University of Illinois College of Medicine Peoria(伊利诺伊大学皮奥里亚医学院儿科部) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机与数据科学学院) Carle Illinois College of Medicine, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校卡莱医学院)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 提出利用图对齐拓扑作为归纳偏置,通过构建参考信息与LLM输出之间的对齐二分图并训练图神经网络建模对齐结构,在幻觉检测和问答任务上取得最先进结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23825 2026-05-25 cs.LG cs.AI 62%

It's the humans, not the data: Geopolitical bias in LLMs originates in post-training, amplified by the language of the prompt

是人类,而非数据:LLM中的地缘政治偏见源于后训练,并通过提示语言放大

Stuart Bladon, Brinnae Bent

机构 * Alibaba(阿里巴巴) seven AI labs(七家人工智能实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 研究发现大语言模型的地缘政治偏见主要源于后训练阶段而非预训练,且偏见方向与模型开发者所在国一致,提示语言会放大该偏见。

Comments 12 pages, 6 figures, 2 tables, 3 appendices. Code and scenario bank: https://github.com/recozers/LLM-Bias

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23058 2026-05-25 cs.SE cs.AI 57%

A measurement substrate for agentic Kubernetes operations: Methodology and a case study in retrieval-compounding falsification

面向代理化 Kubernetes 操作的测量基础:方法论与检索复合证伪案例研究

Joshua Odmark, Gideon Rubin, Deon van der Vyver

机构 * Independent(独立) LDE Cognyx

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 提出一个闭环测量框架 agent-breakage,通过注入故障、评分响应并累积带标签元组,解决 Kubernetes 操作代理经验声明的不可证伪性问题,并在案例研究中发现检索复合能力被三个混杂因素(pgvector 索引错误、+19% 选择偏差、小样本夸大效应约 3 倍)所掩盖。

Comments 22 pages. Code at https://github.com/odmarkj/agent-breakage tag v0.1.0 (Apache 2.0). Source repo at https://github.com/odmarkj/agent-breakage-paper tag arxiv-v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23330 2026-05-25 cs.CR 50%

Security, Privacy, and Ethical Risks in OpenClaw

OpenClaw 中的安全、隐私与伦理风险

Yutong Jin, Zelin Zhang, Zhijin Lyu, Jianbing Ni

专题命中 AI治理与伦理 :trustworthy(abstract)

AI总结 本文系统研究了本地可执行AI代理系统OpenClaw在安全、隐私、伦理及可追溯性方面的风险,并呼吁多方合作构建更安全可靠的AI代理系统。

Comments Accepted by Journal of Information and Intelligence(JII)

详情

展开后加载摘要…

URL PDF HTML 收藏