Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
基于检测器的视频大语言模型用于高效的时空定位
Shida Gao, Feng Xue, Xiangfeng Wang, Anlong Ming, Zhaowen Lin, Haiyang Zhang, Teng Long, Nicu Sebe, Yihua Shao, Haozhe Wang, Wei Wang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of Trento(特伦特大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Hong Kong University of Science and Technology(香港科技大学)
;
ZTE Corporation(中兴通讯)
All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept Reproduction
所有变化可能都有不变原则:通过设计概念再现改进永动有害迷因检测
Ziyou Jiang, Mingyang Li, Junjie Wang, Yuekai Huang, Jie Huang, Zhiyuan Chang, Zhaoyang Li, Qing Wang
机构
*
State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室)
;
Science and Technology on Integrated Information System Laboratory Institute of Software Chinese Academy of Sciences(软件研究所信息集成系统技术研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
机构
*
Institute for AI, Peking University(人工智能研究院,北京大学)
;
Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院)
;
School of Psychological and Cognitive Sciences, Peking University(心理学与认知科学学院,北京大学)
;
School of Computer Science, Peking University(计算机科学学院,北京大学)
;
Yuanpei College, Peking University(元培学院,北京大学)
;
School of Foreign Languages, Peking University(外语学院,北京大学)
;
School of EECS, Peking University(电子工程与科学学院,北京大学)
;
Huazhong University of Science and Technology(华中科技大学)
;
State Key Lab of General AI(通用人工智能国家重点实验室)
;
Nat’l Eng. Research Center of Visual Technology(视觉技术国家工程研究中心)
;
Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京行为与心理健康重点实验室,北京大学)
;
Embodied Intelligence Lab, PKU-Wuhan Institute for Artificial Intelligence(具身智能实验室,北京大学-武汉人工智能研究院)
Bisecle: Binding and Separation in Continual Learning for Video Language Understanding
Yue Tan, Xiaoqian Hu, Hao Xue, Celso De Melo, Flora D. Salim
机构
*
School of Computer Science University of New South Wales(计算机科学学院新南威尔士大学)
;
School of Computer Science and Engineering University of New South Wales(计算机科学与工程学院新南威尔士大学)
;
DEVCOM Army Research Laboratory(陆军研究实验室)
专题命中
视频多模态
:multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV
机构
*
Central South University(中南大学)
;
Tsinghua University(清华大学)
;
South China Normal University(华南师范大学)
;
ByteDance Inc(字节跳动公司)
;
University of Zaragoza(阿拉维达大学)
;
CosmosMind
;
Wuhan University(武汉大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
Southeast University(东南大学)
;
Tencent(腾讯公司)
;
Nankai University(南开大学)
;
Supermicro Computer Inc(Supermicro计算机公司)
;
Huazhong University of Science and Technology(华中科技大学)
MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering
MemoryCard: 面向长视频问答的主题感知多模态线索压缩
Qing Yang, Pengcheng Huang, Xinze Li, Zhenghao Liu, Yukun Yan, Yu Gu, Ge Yu, Gang Li, Maosong Sun
机构
*
School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Digital China Group(数字中国集团)
ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling
ForestPrune: 通过时空森林建模实现视频多模态大语言模型的高比率视觉令牌压缩
Shaobo Ju, Baiyang Song, Tao Chen, Jiapeng Zhang, Qiong Wu, Chao Chang, HuaiXi Wang, Yiyi Zhou, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
;
National University of Defense Technology(国防科技大学)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
Yujia Liang, Jile Jiao, Xuetao Feng, Zixuan Ye, Yuan Wang, Zhicheng Wang
机构
*
School of AIA, Huazhong University of Science and Technology(华中科技大学人工智能学院)
;
Deepeleph Intelligent Technology(DeepEleph智能技术)
;
JD Explore Academy(JD探索学院)
;
Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)