FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction
机构 * AMAP, Alibaba Group(阿里集团AMAP) ; Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 视频理解 :video generation(abstract);分类 cs.CV
AI 大模型
视频理解、视频生成、视频语言模型和时序视觉推理。
机构 * AMAP, Alibaba Group(阿里集团AMAP) ; Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 视频理解 :video generation(abstract);分类 cs.CV
机构 * Leibniz University Hannover(莱比锡大学汉诺威分校) ; L3S Research Center(L3S研究中心) ; Microsoft(微软公司)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments Accepted to NeurIPS 2025
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Zhejiang University(浙江大学) ; The University of Tokyo(东京大学) ; Fudan University(复旦大学) ; Nanjing University(南京大学)
专题命中 视频理解 :video reasoning(abstract);分类 cs.CV
Comments Accepted at NeurIPS 2025
机构 * Southeast University(东南大学) ; Monash University(墨尔本大学) ; Xiaohongshu Inc.(小红书公司) ; University of Southern California(南加州大学) ; Fudan University(复旦大学)
专题命中 视频理解 :video reasoning(abstract);分类 cs.CV
机构 * University of Science and Technology of China(中国科学技术大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Beihang University(北京航空航天大学) ; Shanghai Jiao Tong University(上海交通大学) ; Zhejiang University(浙江大学) ; State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) ; Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
机构 * University of Bristol(布里斯托大学)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments Accepted at eLVM workshop at CVPR 2025
机构 * University of Southern California(南加州大学) ; USC Institute for Creative Technologies(USC创意技术研究所)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments 8 pages, 5 figures
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, USTC(脑启发智能感知与认知国家重点实验室,中国科学技术大学) ; Shanghai Innovation Institute(上海创新研究院) ; Tencent QQ(腾讯QQ) ; Fudan University(复旦大学)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments 34 pages, 25 figures
机构 * North Dakota State University(北达科塔州立大学) ; University of Missouri–Columbia(密苏里大学哥伦比亚分校) ; University of Memphis(孟菲斯大学)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments This paper was accepted at ICCV 2025
机构 * School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学,南京,中国) ; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其跨学科应用关键实验室(东南大学),教育部,中国)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) ; Amazon(亚马逊)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments Accepted at ECCV 2024
机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队) ; New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所模式识别新实验室) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; Peking University(北京大学) ; Nanjing University(南京大学)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments Project webpage: https://avocado-captioner.github.io/
机构 * Nagoya University(名古屋大学)
专题命中 视频理解 :video-language(abstract);分类 cs.CV
Comments ACM Multimedia Asia 2025
机构 * Bytedance(字节跳动) ; Peking University(北京大学) ; Peking University People’s Hospital(北京大学人民医院) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
专题命中 视频理解 :video-language(abstract);分类 cs.CV
Comments Accepted to AAAI 2025
机构 * ECE Department(电子工程系) ; University of California(加州大学)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments 4 pages, 2 figures
机构 * School of Mathematics, Statistics and Computer Science, College of Science, University of Tehran(数学、统计与计算机科学学院,科学学院,塔里斯坦大学)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
机构 * Guilin University of Electronic Technology(桂林电子科技大学) ; The Ohio State University(俄亥俄州立大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; University of California, Berkeley(加州大学伯克利分校)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
机构 * Arizona State University(亚利桑那州立大学) ; Indian Institute of Technology, Kharagpur(印度理工学院,克拉格浦)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments 17 pages, 9 figures, 6 tables. Presents TimeWarp, a synthetic preference data framework to improve temporal understanding in Video-LLMs, showing consistent gains across seven benchmarks. Includes supplementary material in the Appendix
机构 * department of Computer Science, University of Maryland, College Park, MD, 20742(计算机科学系,马里兰大学, College Park, MD)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments Video link: https://drive.google.com/file/d/1UccngwgPqUwPMhBja7JrXfZoTquCx_Qe/view?usp=sharing
机构 * University of Bristol(布里斯托大学) ; Memories.ai Research(Memories.ai研究)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
机构 * Peking University(北京大学) ; AI Geeks ; Australian Artificial Intelligence Institute(澳大利亚人工智能研究所)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
机构 * The ADAPT SFI Research Centre(ADAPT SFI研究机构) ; School of Computer Science, University College Dublin(大学学院计算机科学系) ; Amazon Development Center Germany GmbH(亚马逊德国开发中心) ; Beijing-Dublin International College(北京-都柏林国际学院)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments 6 pages, 3 figures
机构 * Arizona State University(亚利桑那州立大学)
专题命中 视频理解 :video-language(abstract);分类 cs.CV
Comments EMNLP 2025 (Main)
机构 * University of Trento(特伦托大学) ; Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments Accepted to NeurIPS 2025
机构 * Korea University(韩国大学) ; Hanwha Vision(翰威英航) ; KAIST(韩国科学技术院)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments EMNLP 2025 Findings
机构 * Harbin Institute of Technology(哈尔滨理工大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
机构 * Department of Computer Science, Iowa State University, Ames, IA, USA(计算机科学系,爱荷华州立大学) ; Department of Civil, Construction and Environmental Engineering, Iowa State University, Ames, IA, USA(土木、建设与环境工程系,爱荷华州立大学)
专题命中 视频理解 :video-language(abstract);分类 cs.CV
机构 * Visual Geometry Group, University of Oxford(牛津大学视觉几何组) ; CVIT, IIIT Hyderabad(海得拉巴印度理工学院计算机视觉研究所) ; LIGM, École des Ponts ParisTech(巴黎理工学院路易-狄塞尔数学与计算机科学实验室) ; SAI, Shanghai Jiao Tong University(上海交通大学人工智能研究所)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments ICCV 2025. Project Page: https://www.robots.ox.ac.uk/vgg/research/shot-by-shot/