Unified Graph Structured Models for Video Understanding
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
AI 大模型
视频理解、视频生成、视频语言模型和时序视觉推理。
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments ECCV 2020
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments Accepted by WACV2021
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments 10 pages, 2 figures, submitted to ICASSP 2021
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments ECCV 2020
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments The authors have temporarily withdrawn this paper to reassess some of the experimental results
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments accepted by MM2020 sports workshop
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments Accepted to ECCV 2020
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments Accepted by ICCV 2019
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments PhD dissertation. Posting here for easier access and update. Chapter 4, 6 and 7 are collaborative work
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments Accepted paper at 2018 ECCV Youtube8M workshop: https://research.google.com/youtube8m/workshop2018/
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments CVPR 2018
专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV
Comments To appear on CVPR 2017 YouTube-8M Workshop(Rank 3rd out of 650 teams)
SCOUT:面向超长时长自我中心视频推理的自检查与恢复感知工具思维智能体
机构 * Sun Yat-sen University(中山大学) ; Shenzhen Loop Area Institute(深圳河套学院) ; Guangdong OPPO Mobile Telecommunications Corp., Ltd.(广东欧珀移动通信有限公司) ; OPPO AI Center(OPPO人工智能中心) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中 视频理解 :video reasoning(title);video understanding(abstract)
AI总结 本研究提出SCOUT智能体框架,结合自适应策略与UPS-GRPO方法,解决超长自我中心视频推理的错误传播及信用分配问题,在对应基准上达最优性能且短时长场景仍具竞争力。
Comments Accepted by ACM Multimedia 2026 (MM '26)
EgoEverything:一个受人类行为启发的长上下文第一人称视频理解AR环境基准
机构 * New York University(纽约大学) ; Meta Reality Labs
专题命中 视频理解 :video understanding(title,abstract)
AI总结 本文提出EgoEverything基准,通过利用人类注意力信号生成问题,更真实地捕捉自然人类行为,用于评估AR环境中长上下文第一人称视频理解。
面向高效视频大语言模型的sink-token感知剪枝:用于细粒度视频理解
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国高级科学技术研究院) ; Oracle ; University of California, San Diego(加州大学圣地亚哥分校) ; Electronics and Telecommunications Research Institute (ETRI)(电子电信研究院)
专题命中 视频理解 :video understanding(title,abstract)
AI总结 本文提出Sink-Token-aware Pruning方法,通过识别并抑制semantically uninformative tokens,提升细粒度视频理解性能,在多种基准测试中表现优异。
Comments ECCV 2026
VCIFBench:评估视频理解中的复杂指令遵循能力
机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) ; Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University(知识驱动人机智能工程研究中心,吉林大学) ; International Center of Future Science, Jilin University(未来科学国际中心,吉林大学)
专题命中 视频理解 :video understanding(title,abstract)
AI总结 提出VCIFBench基准,通过混合验证流水线评估多模态大模型在视频理解中遵循内容、格式、风格和结构约束的复杂指令能力,实验表明联合约束满足仍具挑战,DPO训练可提升性能。
HiCrew:通过问题感知多智能体协作实现长视频理解的层次推理
机构 * School of Artificial Intelligence, Sun Yat-sen University, China(中山大学人工智能学院)
专题命中 视频理解 :video understanding(title,abstract)
AI总结 本文提出HiCrew框架,通过混合树结构、问题感知描述生成和动态规划层,解决长视频中时空冗余和叙事依赖问题,提升时间与因果推理能力。
机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) ; DEVCOM Army Research Laboratory(国防部陆军研究实验室)
专题命中 视频理解 :video understanding(title,abstract)
机构 * Weldon School of Biomedical Engineering, Purdue University(普渡大学生物医学工程学院) ; Department of Industrial and Operations Engineering, University of Michigan(密歇根大学工业与运作工程系) ; Edward P. Fitts Department of Industrial and Systems Engineering, North Carolina State University(北卡罗来纳州立大学工业与系统工程系) ; School of Nursing, Purdue University(普渡大学护理学院)
专题命中 视频理解 :video-language(title,abstract)
机构 * The University of Hong Kong(香港大学) ; The Chinese University of Hong Kong(香港中文大学) ; Google(谷歌)
专题命中 视频理解 :video understanding(title);long video(abstract)
Comments Accepted to UIST 2025. Project website: https://zhaorunning.github.io/NoteIt/
机构 * Google DeepMind(谷歌DeepMind)
专题命中 视频理解 :video understanding(title,abstract)
专题命中 视频理解 :video understanding(title,abstract)
专题命中 视频理解 :video understanding(title,abstract)
Comments Appears in the proceedings of the 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2018
专题命中 视频理解 :video understanding(title,abstract)
Comments YouTube-8M Workshop submission, 8 pages
PPLLaVA:通过提示引导的视频序列理解
机构 * Peng Cheng Laboratory(鹏城实验室) ; Peking University(北京大学) ; XPeng Inc.(小鹏汽车有限公司) ; Ministry of Industry and Information Technology of the People’s Republic of China, China Electronics Standardization Institute, Beijing, China(中华人民共和国工业和信息化部,中国电子标准化研究院,北京,中国)
专题命中 视频理解 :video understanding(abstract);video reasoning(abstract);long video(abstract);分类 cs.CV
AI总结 PPLLaVA通过提示引导的池化策略实现视频序列高效理解,减少18倍token,提升推理效率并在多种视频理解任务中取得最佳性能。
Comments Accepted to ICLR' 26
VTok: 一种具有解耦空间-时间潜在变量的统一视频分词器
机构 * Johns Hopkins University(约翰斯·霍普金斯大学)
专题命中 视频理解 :video generation(abstract);video understanding(abstract);text-to-video(abstract);分类 cs.CV
AI总结 VTok通过解耦空间和时间潜在变量,实现紧凑且高效的视频分词,提升视频理解和生成任务的性能与效率。
Conan:基于多尺度视觉证据的逐步学习以像侦探一样推理
机构 * State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) ; WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)
专题命中 视频理解 :video understanding(abstract);video reasoning(abstract);long video(abstract);分类 cs.CV
AI总结 Conan通过多阶段渐进冷启动策略和AIR RLVR框架,实现证据基础的多步视频推理,超越基线模型,达到最先进的性能。
专题命中 视频理解 :video understanding(abstract);video-language(abstract);long video(abstract);分类 cs.CV
Comments Accepted by AAAI 2026