Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments First three authors contributed equally. Dataset: https://huggingface.co/datasets/VLLMs/MIRB