Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV
Comments Our dataset can be found at \url{https://huggingface.co/datasets/fazliimam/temporal-vqa}