How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
GPT-4o对视觉的理解有多好?在标准计算机视觉任务上评估多模态基础模型
机构 * Swiss Federal Institute of Technology(瑞士联邦理工学院)
专题命中 多模态生成 :multimodal(title,abstract);multimodal foundation model(title,abstract);image-text(abstract);分类 cs.CV、cs.AI
AI总结 本文评估了GPT-4o等多模态基础模型在标准计算机视觉任务上的表现,发现其在语义任务上优于几何任务,GPT-4o在非推理模型中表现最佳,推理模型在几何任务中有所提升。
Comments ICLR 2026. Project page at https://fm-vision-evals.epfl.ch/