Video Understanding by Design: How Datasets Shape Video Models
通过设计理解视频:数据集如何塑造视频模型
Lei Wang, Syuan-Hao Li, Piotr Koniusz, Yongsheng Gao
机构
*
School of Engineering and Built Environment, Electrical and Electronic Engineering, Griffith University(工程与建筑环境学院,电气与电子工程学院,格里菲斯大学)
;
School of Computer Science and Engineering, University of New South Wales(计算机科学与工程学院,新南威尔士大学)
专题命中
视频多模态
:multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI
Toward Scalable Co-located Practical Learning: Assisting with Computer Vision and Multimodal Analytics
迈向可扩展的协同实践学习:协助计算机视觉和多模态分析
Xinyu Li, Linxuan Zhao, Yueqiao Jin, Yuchen Liu, Jin Zhou, Roberto Martinez-Maldonado, Dragan Gasevic, Lixiang Yan
机构
*
Centre for Learning Analytics at Monash(墨尔本大学学习分析中心)
;
Monash University(墨尔本大学)
;
Department of Civil and Environmental Engineering(土木与环境工程系)
;
School of Education(教育学院)
;
The University of Hong Kong(香港大学)
Decoupling Semantics and Logic: A Training-Free Coarse-to-Fine Pipeline for Video Retrieval-Augmented Generation
解耦语义与逻辑:一种无需训练的从粗到精的视频检索增强生成流水线
Jiaxin Dai, Zehang Wei, Jiamin Yan, Xiang Xiang
机构
*
School of Computer Science & Tech, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)
;
School of AI and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
Self-supervised Learning Matters: A Simple Ensemble Solution for Micro-Gesture Recognition
自监督学习至关重要:一种用于微手势识别的简单集成方案
Tingyi Liu, Kun Li, Fei Wang, Junjie Chen, Zhiliang Wu, Jihao Gu, Haixu Liu, Dan Guo
机构
*
Hefei University of Technology(合肥工业大学)
;
United Arab Emirates University(阿拉伯联合酋长国大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
;
Anhui Evolution Technology Co., Ltd.(安徽进化科技有限公司)
;
Nanyang Technological University(南洋理工大学)
;
University College London(伦敦大学学院)
;
The University of Sydney(悉尼大学)
;
Beijing QBoson Quantum Technology Co., Ltd.(北京量子芯科技有限公司)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
;
Harbin Institute of Technology (Weihai)(哈尔滨工业大学(威海))
Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation
Dream-Tac: 用于接触丰富机器人操作任务的统一触觉世界动作模型
Yunfan Lou, Yifan Ye, Yankai Fu, Jun Cen, Xiaowei Chi, Yaoxu Lyu, Peidong Jia, Sirui Han, Zhihe Lu, Shanghang Zhang
机构
*
Peking University(北京大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Nanjing University(南京大学)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室)
MM-Matryoshka: Towards Budget-Elastic Visual Document Retrieval via a 2D Multimodal Matryoshka Training Framework
MM-Matryoshka:通过二维多模态套娃训练框架实现预算弹性视觉文档检索
Haowen Xiang, Yibo Yan, Jiahao Huo, Yu Huang, Yi Cao, Mingdong Ou, Xuming Hu
机构
*
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Alibaba Cloud Computing(阿里云计算)
;
Hong Kong University of Science and Technology(香港科技大学)
CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval
CR-JEPA:用于遥感图像检索的跨模态联合嵌入预测学习
Md Aminur Hossain, Ayush V. Patel, Nitant Dube, Biplab Banerjee
机构
*
Space Applications Centre, Indian Space Research Organisation(印度空间研究组织空间应用中心)
;
Centre of Studies in Resources Engineering, Indian Institute of Technology Bombay(印度理工学院孟买资源工程研究中心)