Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
在视觉-语言模型中可靠性在哪里存在:注意力、隐藏状态和因果回路的机制研究
机构 * UC Santa Barbara(加州大学圣巴巴拉分校) ; UC Berkeley(加州大学伯克利分校) ; NVIDIA(英伟达) ; Algoverse AI Research(Algoverse人工智能研究) ; Brown University(布朗大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
AI总结 本文通过机制性研究发现,视觉-语言模型的可靠性主要体现在隐藏状态几何、分层边际形成和稀疏晚层回路,而非注意力图的锐度。
Comments 15 pages, 4 figures, 10 tables. Accepted at the ICLR 2026 Workshop on Multimodal Reasoning. Code and probe-training pipelines: https://github.com/itsloganmann/VLM-Reliability-Probe