From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability
从恢复到骤降:动作后训练如何降低视觉语言模型的深层解码能力
机构 * New York University(纽约大学) ; Reflex ; Université de Montréal(蒙特利尔大学) ; Maastricht University(马斯特里赫特大学)
专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV
AI总结 该研究探究动作后训练对VLM空间理解的影响,发现VLA存在深层解码能力的“基底”差距与“悬崖”骤降,定位其源于深层MLP干扰,消融该模块可恢复多数解码性能。
Comments Accepted to the archival proceedings track of the Embodied Multimodal Reasoning (EMR) Workshop at ECCV 2026