Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
通过跨模态注意力可区分性理解视频-语言模型中的时间逻辑一致性
机构 * School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China(北京理工大学计算机科学与技术学院,北京,中国) ; Beijing Engineering Research Center of High Volume Language Information Processing and Cloud Computing Applications, Beijing Institute of Technology, Beijing, China(高性能语言信息处理与云计算应用北京工程研究中心,北京理工大学,北京,中国)
专题命中 其他视频模型 :video-language(title,abstract);分类 cs.CV、cs.MM
AI总结 研究探讨视频-语言模型在时间逻辑一致性问题上的核心原因,提出TCAS方法提升跨模态注意力的时序分辨能力,实验验证了方法对时间逻辑一致性的提升效果。
Comments Accepted by CVPR 2026