Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering
用于通用文档视觉问答的领域适应视觉语言模型的比较研究
机构 * Universidad Autónoma de Madrid (UAM)(马德里自治大学) ; BiometricsAI(生物识别人工智能)
专题命中 视觉问答 :visual question answering(title,abstract);vision-language model(abstract);分类 cs.CV、cs.LG
AI总结 研究对8个开源预训练VLMs在三种文档领域的DocVQA进行全面评估,通过多种评估方式发现其在不同布局性能有差异,参数缩放影响性能,视觉理解是瓶颈,还表明少样本微调能让模型快速适应目标域文档。
Comments 17 pages, 4 figures, accepted at the Automatically Domain-Adapted and Personalized Document Analysis workshop of the ICDAR 2026