Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
机构 * Institute of Automation, CAS(中国科学院自动化研究所) ; School of Artifcial Intelligence, UCAS(中国科学技术大学人工智能学院) ; Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院) ; Shenzhen International GraduateSchool,Tsinghua University(深圳国际研究生院,清华大学) ; State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,北京大学计算机学院)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV