arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-05-15 至 2026-05-15 共收录 66 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 16 篇

2605.14535 2026-05-15 cs.LG 57%

Exploring Geographic Relative Space in Large Language Models through Activation Patching

通过激活修补探索大语言模型中的地理相对空间

Stef De Sabbata, Rahul Baiju, Stefano Mizzaro, Kevin Roitero

机构 * School of Geography, Geology and the Environment, University of Leicester, UK(地理、地质与环境学院,莱斯特大学,英国) Department of Mathematics, Computer Science and Physics, University of Udine, Italy(数学、计算机科学与物理系,乌迪内大学,意大利)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 研究大语言模型处理地理相对空间的机制,通过激活修补技术揭示其内部运作,提升地理应用的安全性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14025 2026-05-15 q-bio.NC cs.AI 57%

Do Language Models Align with Brains? Prediction Scores Are Not Enough

语言模型与大脑对齐吗?预测分数并不足够

Xiao Jia

机构 * School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(人工智能学院,香港中文大学(深圳))

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文通过L-PACT框架评估语言模型与大脑的对齐性,发现预测分数不足以支持两者对齐的结论,所有测试结果均被控制条件解释。

Comments 39 pages, 4 main figures, 6 supplementary figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15054 2026-05-15 cs.CV 50%

LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection

LATERN:测试时上下文感知的可解释视频异常检测

Mitchell Piehl, Muchao Ye

机构 * The University of Iowa(爱荷华大学)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出LATERN框架,通过上下文感知模块和递归证据聚合模块提升视频异常检测的准确性和解释一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14705 2026-05-15 cs.CV 50%

Towards Continuous Sign Language Conversation from Isolated Signs

从孤立手势构建连续手语对话

Youngmin Kim, Kyobin Choo, Jiwoo Park, Minseo Kim, Chanyoung Kim, Junhyeok Kim, Seong Jae Hwang

机构 * Yonsei University(延世大学) LG Electronics(LG电子) Emory University(埃默里大学)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出通过构建连续手语对话来解决手语翻译模型词汇覆盖不足和泛化能力弱的问题,引入了SignaVox-W和SignaVox-U数据集,并采用检索引导的语音到词组翻译和BRAID模型实现结构匹配。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14675 2026-05-15 cs.SE 50%

Agentic AI in Industry: Adoption Level and Deployment Barriers

工业中的代理AI:采用级别与部署障碍

Spyridon Alvanakis Apostolou, Jan Bosch, Helena Holmström Olsson

专题命中 其他安全 :safety(abstract)

AI总结 研究通过16次访谈,揭示工业组织代理AI的采用现状及部署障碍,发现能力与部署验证存在差距,主要障碍包括LLM上下文窗口限制、专有语言性能不足、非确定性与数据保密问题。

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02271 2026-05-15 cs.CV 50%

Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework

医学报告生成:基于层次任务结构的跨模态因果干预框架

Yucheng Song, Yifan Ge, Junhao Li, Zhining Liao, Zhifang Liao

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出HTSC-CIF框架,通过层次任务分解解决医学报告生成中领域知识不足、跨模态对齐差和虚假相关性三大问题,提升生成效果。

Comments Due to issues with the training epochs and training strategy in our paper, there are numerical errors in the result comparison table presented in the preprint. Therefore, we have decided to withdraw the manuscript for further revision

详情

展开后加载摘要…

URL PDF HTML 收藏