arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-05-08 至 2026-05-08 共收录 4 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 4 篇

2605.06537 2026-05-08 cs.CV 83%

MedHorizon: Towards Long-context Medical Video Understanding in the Wild

MedHorizon:迈向真实场景下的长上下文医学视频理解

Bodong Du, Bowen Liu, Yang Yu, Xinpeng Ding, Zhiheng Wu, Shuning Wang, Shuo Nie, Naiming Liu, Qifeng Chen, Yangqiu Song, Xiaomeng Li

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Baidu Inc.(百度公司) Xidian University(西安电子科技大学)

专题命中 视频理解 :video understanding(title,abstract);long video(abstract);分类 cs.CV

AI总结 MedHorizon提出一个真实场景下的长上下文医学视频理解基准,通过稀疏证据和多跳临床推理评估,揭示当前系统在完整流程理解上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14724 2026-05-08 cs.CV cs.AI cs.CL 79%

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding

HERMES: KV缓存作为分层内存用于高效视频流理解

Haowei Zhang, Shudong Yang, Jinlan Fu, See-Kiong Ng, Xipeng Qiu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) National University of Singapore(新加坡国立大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 HERMES提出一种无需训练的架构,通过将KV缓存视为分层内存框架,实现视频流的实时准确理解,提升处理速度并降低内存消耗。

Comments Accepted to ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13281 2026-05-08 cs.CV 70%

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?

VideoASMR-Bench: AI生成的ASMR视频能否欺骗视觉语言模型和人类?

Jiaqi Wang, Weijia Wu, Yi Zhan, Rui Zhao, Ming Hu, James Cheng, Wei Liu, Philip Torr, Kevin Qinghong Lin

机构 * The Chinese University of Hong Kong(香港中文大学) National University of Singapore(新加坡国立大学) Peking University(北京大学) Monash University(墨尔本大学) Video Rebirth University of Oxford(牛津大学)

专题命中 视频理解 :video generation(abstract);video understanding(abstract);分类 cs.CV

AI总结 VideoASMR-Bench通过细粒度音频视觉感知和感官沉浸性评估AI生成ASMR视频的检测能力,揭示了当前VLMs在识别AI生成ASMR视频上的不足以及VGMs生成逼真ASMR视频的能力。

Comments Code is at https://github.com/video-reality-test/video-reality-test, page is at https://video-reality-test.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24382 2026-05-08 cs.CV cs.AI 57%

REMAP: Regularized Matching and Partial Alignment of Video Embeddings

REMAP:基于正则化的视频嵌入匹配与部分对齐

Soumyadeep Chandra, Kaushik Roy

机构 * Elmore Family School of Electrical and Computer Engineering(埃尔摩家庭电气与计算机工程学院)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 REMAP通过正则化的融合部分Gromov-Wasserstein最优传输框架,解决现实视频中噪声、冗余和执行差异等问题,提升过程学习效果。

Comments 9 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏