arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-04-24 至 2026-04-24 共收录 10 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 3 篇

2604.09000 2026-04-24 cs.CV 79%

StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding

StreamMeCo:用于高效流视频理解的长期智能体内存压缩

Junxi Wang, Te Sun, Jiayi Zhu, Junxian Li, Haowen Xu, Zichen Wen, Xuming Hu, Zhiyu Li, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) Hong Kong University of Science and Technology(香港科学与技术大学) MemTensor (Shanghai) Technology Co., Ltd.(MemTensor(上海)技术有限公司)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 StreamMeCo通过内存图连通性引入边自由minmax采样和边感知权重剪枝,有效压缩内存并保持精度,实验表明在70%压缩下实现1.87倍检索速度提升和1.0%的准确率提升。

Comments 2026ACL Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21444 2026-04-24 cs.AI 78%

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration

HiCrew:通过问题感知多智能体协作实现长视频理解的层次推理

Yuehan Zhu, Jingqi Zhao, Jiawen Zhao, Xudong Mao, Baoquan Zhao

机构 * School of Artificial Intelligence, Sun Yat-sen University, China(中山大学人工智能学院)

专题命中 视频理解 :video understanding(title,abstract)

AI总结 本文提出HiCrew框架,通过混合树结构、问题感知描述生成和动态规划层,解决长视频中时空冗余和叙事依赖问题,提升时间与因果推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21011 2026-04-24 cs.CV q-bio.NC 57%

Micro-DualNet: Dual-Path Spatio-Temporal Network for Micro-Action Recognition

微双网:用于微动作识别的双路径时空网络

Naga VS Raviteja Chappa, Evangelos Sariyanidi, Lisa Yankowitz, Gokul Nair, Casey J. Zampella, Robert T. Schultz, Birkan Tunç

机构 * The Children’s Hospital of Philadelphia(费城儿童医院) University of Pennsylvania(宾夕法尼亚大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 本文提出双路径网络,通过并行的空间-时间(ST)和时间-空间(TS)路径处理人体部位,采用自适应路由和MAC损失提升微动作识别性能。

Comments Accepted to International Conference on Automatic Face and Gesture Recognition (FG)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 4 篇

2604.21221 2026-04-24 cs.CV cs.LG 85%

Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation

稀疏强制:用于实时自回归扩散视频生成的原生可训练稀疏注意力

Boxun Xu, Yuming Du, Zichang Liu, Siyu Yang, Ziyang Jiang, Siqi Yan, Rajasi Saha, Albert Pumarola, Wenchen Wang, Peng Li

机构 * Meta Superintelligence Labs(Meta超智能实验室) University of California, Santa Barbara(加州大学圣芭芭拉分校)

专题命中 视频生成 :video generation(title,abstract);video diffusion(abstract);text-to-video(abstract);分类 cs.CV

AI总结 本文提出稀疏强制方法,通过在训练和推理中提升长时程生成质量并降低解码延迟,采用可训练的原生稀疏机制和高效的GPU内核PBSA来加速稀疏注意力和内存更新。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21362 2026-04-24 cs.CV 83%

KD-CVG: A Knowledge-Driven Approach for Creative Video Generation

KD-CVG:一种基于知识的创意视频生成方法

Linkai Liu, Wei Feng, Xi Zhao, Shen Zhang, Xingye Chen, Zheng Zhang, Jingjing Lv, Junjie Shen, Ching Law, Yuchen Zhou, Zipeng Guo, Chao Gou

机构 * Sun Yat-sen University(中山大学)

专题命中 视频生成 :video generation(title,abstract);text-to-video(abstract);分类 cs.CV

AI总结 本文提出KD-CVG方法,通过构建广告创意知识库解决文本到视频模型在语义对齐和运动适应性方面的不足,提升创意视频生成效果。

Comments Accepted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21291 2026-04-24 cs.CV cs.AI 79%

Exploring the Role of Synthetic Data Augmentation in Controllable Human-Centric Video Generation

探索合成数据增强在可控人类中心视频生成中的作用

Yuanchen Fei, Yude Zou, Zejian Kang, Ming Li, Jiaying Zhou, Xiangru Huang

机构 * Hunan University(湖南大学) Westlake University(西湖大学) Shanghai Jiaotong University(上海交通大学) Zhejiang University(浙江大学) Shanghai Innovation Institute(上海创新研究院) Sun Yat-Sen University(中山大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

AI总结 本文研究合成数据对可控人类视频生成的影响,提出基于扩散的框架实现细粒度控制,并通过实验揭示合成与真实数据的互补作用,为高效选择合成样本提升视频真实感提供方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07940 2026-04-24 cs.CV 57%

Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos

超越帧限:从视角视频生成360度全景视频

Rundong Luo, Matthew Wallingford, Ali Farhadi, Noah Snavely, Wei-Chiu Ma

机构 * Cornell University(康奈尔大学) University of Washington(华盛顿大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

AI总结 本文提出从视角视频生成完整全景视频的方法,通过设计几何与运动感知操作提升生成质量,实现空间与时间一致性,应用于视频稳定、视角控制和交互式视觉问答。

Comments Project page: https://red-fairy.github.io/argus/

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 1 篇

2604.20936 2026-04-24 cs.MM cs.CV cs.HC 84%

AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe

AttentionBender: 在视频扩散变换器中操纵交叉注意力作为创造性探测工具

Adam Cole, Mick Grierson

机构 * University of the Arts London(伦敦艺术大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV、cs.MM

AI总结 AttentionBender通过操纵视频扩散变换器中的交叉注意力,帮助艺术家探索黑箱视频生成的内部机制,揭示交叉注意力的高度交织特性,并产生新的美学效果。

Comments To appear in the Proceedings of the 2026 ACM Creativity and Cognition (C&C '26). 15 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 长视频与时序推理 2 篇

2604.20311 2026-04-24 cs.MM cs.AI 57%

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction

看得更远且更广:面向微视频流行度预测的联合时空放大

Dali Wang, Yunyao Zhang, Junqing Yu, Yi-Ping Phoebe Chen, Chen Xu, Zikai Song

机构 * Huazhong University of Science and Technology(华中科技大学) La Trobe University(拉特罗布大学) Beijing Institute of Computer Technology and Applications(北京计算机技术与应用研究所)

专题命中 长视频与时序推理 :long video(abstract);分类 cs.MM

AI总结 本文提出联合时空放大框架,通过时空增强模块提升微视频流行度预测的精度与可扩展性,实验表明其在主流指标上优于11个基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21154 2026-04-24 cs.AI 50%

Agentic AI for Personalized Physiotherapy: A Multi-Agent Framework for Generative Video Training and Real-Time Pose Correction

代理AI用于个性化物理治疗:一种多代理框架用于生成视频训练和实时姿态校正

Abhishek Dharmaratnakar, Srivaths Ranganathan, Anushree Sinha, Debanshu Das

专题命中 长视频与时序推理 :video generation(abstract)

AI总结 本文提出多代理系统框架,结合生成AI和计算机视觉,解决家庭物理治疗依从性低的问题,通过生成个性化视频和实时姿态校正提升康复效果。

Comments 3 pages, 2 figures, submitted to ICDH IEEE conference

详情

展开后加载摘要…

URL PDF HTML 收藏