arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-03-02 至 2026-03-02 共收录 2 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 2 篇

2602.22745 2026-03-02 cs.CV 79%

SPATIALALIGN: Aligning Dynamic Spatial Relationships in Video Generation

SPATIALALIGN: 视频生成中动态空间关系的对齐

Fengming Liu, Tat-Jen Cham, Chuanxia Zheng

机构 * College of Computing(计算学院) Data Science, Nanyang Technological University, 50 Nanyang Avenue, Singapore 639798(数据科学,南洋理工大学,50 Nanyang Avenue,新加坡639798)

专题命中 视频生成 :video generation(title);text-to-video(abstract);分类 cs.CV

AI总结 SPATIALALIGN通过引入DSR-SCORE和DPO方法,提升文本到视频生成中动态空间关系的对齐能力。

Comments Project page: https://fengming001ntu.github.io/SpatialAlign/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23836 2026-03-02 cs.CV 57%

Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts

立体说话者:基于优先级的混合专家引导的音频驱动3D人类合成

Xiang Deng, Youxin Pang, Xiaochen Zhao, Chao Xu, Lizhen Wang, Hongjiang Xiao, Shi Yan, Hongwen Zhang, Yebin Liu

机构 * Department of Automation, Tsinghua University(自动化系,清华大学) Bytedance Inc.(字节跳动公司) State Key Laboratory of Media Convergence and Communication, Communication University of China(媒体融合与传播国家重点实验室,中国传媒大学) School of Artificial Intelligence, Beijing Normal University(人工智能学院,北京师范大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

AI总结 Stereo-Talker通过优先级引导的混合专家机制实现音频驱动的3D人类合成,生成具有精确唇同步和逼真质量的视频。

Journal ref TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏