CommentsPenguin-VL demonstrates that text-only initialized vision encoders can achieve superior performance in multimodal understanding tasks; Code: https://github.com/tencent-ailab/Penguin-VL
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
Lumos-1:从统一模型视角看基于离散扩散的自回归视频生成
Hangjie Yuan, Weihua Chen, Jun Cen, Hu Yu, Jingyun Liang, Shuning Chang, Zhihui Lin, Tao Feng, Pengwei Liu, Jiazheng Xing, Hao Luo, Jiasheng Tang, Fan Wang, Yi Yang
机构
*
Texas A&M University(德克萨斯A&M大学)
;
University of Minnesota(明尼苏达大学)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
University of Texas at Austin(得克萨斯大学奥斯汀分校)
;
Amazon(亚马逊)
;
State University of New York at Stony Brook(纽约州立大学石溪分校)
AnyCrowd: Instance-Isolated Identity-Pose Binding for Arbitrary Multi-Character Animation
AnyCrowd:实例隔离的身份-姿态绑定用于任意多角色动画
Zhenyu Xie, Ji Xia, Michael Kampffmeyer, Panwen Hu, Zehua Ma, Yujian Zheng, Jing Wang, Zheng Chong, Xujie Zhang, Xianhang Cheng, Xiaodan Liang, Hao Li
机构
*
Mohamed bin Zayed University of Artificial Intelligence, UAE(穆罕默德·本·扎耶德人工智能大学,阿联酋)
;
University of Tromsø (UiT) – The Arctic University of Norway, Norway(特罗姆瑟大学(UiT)——挪威北极大学,挪威)
;
Shenzhen campus of Sun Yet-sen University, China(孙中山大学深圳校区,中国)