arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-03-06 至 2026-03-06 共收录 2 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2603.04977 2026-03-06 cs.CV 89%

Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding

思考,然后验证:一种用于长视频理解的假设-验证多智能体框架

Zheng Wang, Haoran Chen, Haoxuan Qin, Zhipeng Wei, Tianwen Qian, Cong Bai

机构 * College of Computer Science, Zhejiang University of Technology(浙江工业大学计算机科学学院) Zhejiang Key Laboratory of Visual Information Intelligent Processing(浙江视觉信息智能处理重点实验室) UC Berkeley(伯克利大学) College of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);video reasoning(abstract);分类 cs.CV

AI总结 VideoHV-Agent通过结构化假设-验证流程提升长视频理解的准确性、可解释性和效率。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00141 2026-03-06 cs.CV cs.AI 88%

FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding

FLoC:基于设施定位的高效视觉令牌压缩用于长视频理解

Janghoon Cho, Jungsoo Lee, Munawar Hayat, Kyuwoong Hwang, Fatih Porikli, Sungha Choi

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 FLoC提出了一种基于设施定位函数的高效视觉令牌压缩方法,通过懒惰贪心算法实现紧凑且多样化的令牌选择,有效提升长视频理解的处理效率和性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏