Video Understanding: From Geometry and Semantics to Unified Models
视频理解:从几何与语义到统一模型
Zhaochong An, Zirui Li, Mingqiao Ye, Feng Qiao, Jiaang Li, Zongwei Wu, Vishal Thengane, Chengzu Li, Lei Li, Luc Van Gool, Guolei Sun, Serge Belongie
机构
*
Department of Computer Science(计算机科学系)
;
University of Copenhagen(哥本哈根大学)
;
College of Computer Science(计算机科学学院)
;
Nankai University(南开大学)
;
School of Computer and Communication Sciences(计算机与通信科学学校)
;
EPFL(苏黎世联邦理工学院)
;
Department of Computer Science & Engineering(计算机科学与工程系)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
Computer Vision Lab(计算机视觉实验室)
;
University of Würzburg(乌尔姆大学)
;
Computer Science Research Centre(计算机科学研究中心)
;
University of Surrey(萨里大学)
;
School of Electrical, Computer and Telecommunications Engineering(电气、计算机和电信工程学院)
;
University of Wollongong(沃林根大学)
;
Language Technology Lab(语言技术实验室)
;
University of Cambridge(剑桥大学)
;
School of Artificial Intelligence(人工智能学院)
;
Beijing Institute of Technology(北京理工大学)
;
Institute for Computer Science(计算机科学研究所)
;
INSAIT
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
Symphony:一种受认知启发的多智能体系统用于长视频理解
Haiyang Yan, Hongyun Zhou, Peng Xu, Xiaoxue Feng, Mengyi Liu
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Kuaishou Technology(快手科技)
;
School of Future Technology, University of Chinese Academy of Sciences(中国科学院大学未来技术学院)
VideoAtlas: Navigating Long-Form Video in Logarithmic Compute
VideoAtlas: 在对数计算中导航长视频
Mohamed Eltahir, Ali Habibullah, Yazan Alshoibi, Lama Ayash, Tanveer Hussain, Naeemullah Khan
机构
*
King Abdullah University of Science and Technology (KAUST)(卡斯特大学)
;
Department of Computer Science, King Khalid University (KKU)(国王 Khalid 大学计算机科学系)
;
Department of Computer Science, Edge Hill University(Edge Hill 大学计算机科学系)
M2P: Improving Visual Foundation Models with Mask-to-Point Weakly-Supervised Learning for Dense Point Tracking
M2P:通过Mask-to-Point弱监督学习改进视觉基础模型以实现密集点跟踪
Qiangqiang Wu, Tianyu Yang, Bo Fang, Jia Wan, Matias Di Martino, Guillermo Sapiro, Antoni B. Chan
机构
*
Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)
;
Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)
;
Meituan, Shenzhen, China(美团(深圳,中国))
;
School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)
;
Department of Electrical and Computer Engineering, Duke University(杜克大学电气与计算机工程系)
机构
*
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Tencent YouTu Lab(腾讯优图实验室)
;
School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院)
;
Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究重点实验室)
;
Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(教育部下一代智能搜索与推荐工程研究中心)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
MosaicMem: 用于可控视频世界模型的混合空间记忆
Wei Yu, Runjia Qian, Yumeng Li, Liquan Wang, Songheng Yin, Sri Siddarth Chakaravarthy P, Dennis Anthony, Yang Ye, Yidi Li, Weiwei Wan, Animesh Garg
机构
*
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
The University of Osaka(大阪大学)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Mujin Inc.(Mujin公司)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
Taiyuan University of Technology(太原本科技大学)
CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization
CineSRD:利用视觉、听觉和语言线索进行开放世界视觉媒体说话人聚类
Liangbin Huang, Xiaohua Liao, Chaoqun Cui, Shijing Wang, Zhaolong Huang, Yanlong Du, Wenji Mao
机构
*
Hujing Digital Media and Entertainment Group(华景数字媒体与娱乐集团)
;
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS部门)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)
HyperMotionX: The Dataset and Benchmark with DiT-Based Pose-Guided Human Image Animation of Complex Motions
HyperMotionX:基于DiT的Pose引导人体图像动画的Dataset和Benchmark
Shuolin Xu, Siming Zheng, Ziyi Wang, HC Yu, Jinwei Chen, Huaqi Zhang, Daquan Zhou, Tong-Yee Lee, Bo Li, Peng-Tao Jiang
机构
*
Bournemouth University(伯恩茅斯大学)
;
vivo BlueImage Lab, vivo Mobile Communication Co., Ltd(vivo 蓝影实验室,vivo 通信有限公司)
;
Peking University(北京大学)
;
National Cheng-Kung University(国立成功大学)