arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-03-10 至 2026-03-10 共收录 7 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 7 篇

2603.08028 2026-03-10 cs.CV cs.MM 84%

Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades

通过文本到骨骼级联实现可控的复杂人类运动视频生成

Ashkan Taghipour, Morteza Ghahremani, Zinuo Li, Hamid Laga, Farid Boussaid, Mohammed Bennamoun

机构 * Department of Computer Science and Software Engineering, The University of Western Australia(计算机科学与软件工程系,西澳大学) Munich Center for Machine Learning (MCML) and Technical University of Munich (TUM)(慕尼黑机器学习中心(MCML)和技术大学慕尼黑(TUM)) School of Information Technology, Murdoch University(信息科技学院,墨尔本大学) Department of Electrical, Electronics and Computer Engineering, The University of Western Australia(电子、电子与计算机工程系,西澳大学)

专题命中 视频生成 :video generation(title,abstract);video diffusion(abstract);分类 cs.CV、cs.MM

AI总结 本文提出通过文本到骨骼级联框架生成可控复杂人类运动视频,解决文本条件模糊和姿态控制成本高的问题,引入合成数据集并展示模型在多个指标上的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00586 2026-03-10 cs.CV 79%

WildActor: Unconstrained Identity-Preserving Video Generation

WildActor: 无约束身份保持视频生成

Qin Guo, Tianyu Yang, Xuanhua He, Fei Shen, Yong Zhang, Zhuoliang Kang, Xiaoming Wei, Dan Xu

机构 * The Hong Kong University of Science(香港科学与技术大学) National University of Singapore(新加坡国立大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

AI总结 WildActor通过不对称注意力机制和视点自适应采样策略,在无约束视角下实现身份保持的人类视频生成。

Comments Project Page: https://wildactor.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07028 2026-03-10 cs.CR cs.CV 79%

Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking

两个框架至关重要:一种针对文本到视频模型劫持的时序攻击

Moyang Chen, Zonghao Ying, Wenzhuo Xu, Quancheng Zou, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang

机构 * College of Science, Mathematics and Technology, Wenzhou-Kean University(科学、数学与技术学院,温州市凯恩大学) State Key Laboratory of Complex & Critical Software Environment, Beihang University(复杂与关键软件环境国家重点实验室,北航) AI Security Lab(360人工智能安全实验室)

专题命中 视频生成 :text-to-video(title,abstract);分类 cs.CV

AI总结 本文提出TFM攻击方法,通过碎片化提示和隐式替换,提高文本到视频模型的劫持效果,显著提升攻击成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12719 2026-03-10 cs.CV 79%

S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation

S2DiT:用于移动流媒体视频生成的 Sandwich Diffusion Transformer

Lin Zhao, Yushu Wu, Aleksei Lebedev, Dishani Lahiri, Meng Dong, Arpit Sahni, Michael Vasilkovsky, Hao Chen, Ju Hu, Aliaksandr Siarohin, Sergey Tulyakov, Yanzhi Wang, Anil Kag, Yanyu Li

机构 * Snap Inc. Northeastern University(东北大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

AI总结 S2DiT 通过高效注意力机制和 2-in-1 深度学习框架,在移动设备上实现高质量、高速度的流式视频生成。

Comments https://snap-research.github.io/S2DiT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08648 2026-03-10 cs.CV 57%

CAST: Modeling Visual State Transitions for Consistent Video Retrieval

CAST:为一致视频检索建模视觉状态转换

Yanqing Liu, Yingcheng Liu, Fanghong Dong, Budianto Budianto, Cihang Xie, Yan Jiao

机构 * University of California, Santa Cruz(加州大学圣克ruz分校) Google(谷歌)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

AI总结 CAST通过建模视觉状态转换,提升视频检索的一致性和连贯性,适用于多种视觉-语言嵌入空间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07192 2026-03-10 cs.CV 57%

FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis

FastSTAR: 空间时间令牌修剪用于高效的自回归视频合成

Sungwoong Yune, Suheon Jeong, Joo-Young Kim

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

AI总结 FastSTAR通过时空令牌修剪和部分更新机制,实现高效视频生成,在保持质量的同时提升处理速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00481 2026-03-10 cs.HC 50%

From Performers to Creators: Understanding Retired Women's Perceptions of Technology-Enhanced Dance Performance

从表演者到创作者:理解退休女性对技术增强舞蹈表演的认知

Danlin Zheng, Xiaoying Wei, Chao Liu, Quanyu Zhang, Jingling Zhang, Shihui Guo, Mingming Fan

专题命中 视频生成 :video generation(abstract)

AI总结 本文通过交互式舞蹈技术与AI生成内容,帮助退休女性克服年龄相关限制,提升舞台表现并成为表演共创者。

详情

展开后加载摘要…

URL PDF HTML 收藏