Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models
相似性并非逻辑:双编码器视觉语言模型的因式推理
专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV、cs.LG;VLM(comments)
AI总结 研究双编码器视觉语言模型组合约束失效问题,提出因式推理,将证据提取与约束执行分离,引入LCSE方法,在FACTOR-Bench上实验,提升了准确率并保持检索性能。
Comments Accepted at ICML 2026. 18 pages, 8 figures. Project page: https://sultanmo.github.io/factored-vlm Code and benchmark: https://github.com/SultanMo/factored-vlm