VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.CV
Comments COLM 2025. VisOnlyQA dataset, code, and model responses are provided at https://github.com/psunlpgroup/VisOnlyQA. Please also refer to our project website at https://visonlyqa.github.io/