ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling
ForestPrune: 通过时空森林建模实现视频多模态大语言模型的高比率视觉令牌压缩
Shaobo Ju, Baiyang Song, Tao Chen, Jiapeng Zhang, Qiong Wu, Chao Chang, HuaiXi Wang, Yiyi Zhou, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
;
National University of Defense Technology(国防科技大学)
TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis
TESSERA:表面光谱的时间嵌入用于地球表示与分析
Zhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic, Silja Sormunen, Robin Young, Madeline C. Lisaius, Markus Immitzer, Toby Jackson, James Ball, David A. Coomes, Anil Madhavapeddy, Andrew Blake, Srinivasan Keshav
机构
*
University of Cambridge(剑桥大学)
;
dClimate Labs(dClimate实验室)
;
Aalto University(阿尔托大学)
;
University of Bristol(布里斯托大学)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
视觉后期分块:一种上下文分块用于高效视觉文档检索的实证研究
Yibo Yan, Mingdong Ou, Yi Cao, Jiahao Huo, Xin Zou, Shuliang Liu, James Kwok, Xuming Hu
机构
*
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Alibaba Cloud Computing(阿里云)
;
Hong Kong University of Science and Technology(香港科技大学)
机构
*
Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子及计算机工程学系)
;
Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程学系)
;
Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学医学智能与扩展现实研究所)
;
Center for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人创新中心)