Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models
多模态大语言模型中基于谱演化引导的令牌剪枝
Bin Chen, Yuxiang Cai, Yadan Luo, Yi Zhang, Jianwei Yin, Zhi Chen
机构
*
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
Zhejiang Key Laboratory of Digital-Intelligence Service Technology(浙江省数字化服务技术重点实验室)
;
The University of Queensland(昆士兰大学)
;
Singapore Management University(新加坡管理大学)
;
The University of Southern Queensland(南昆士兰大学)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);MLLM(abstract_cn);分类 cs.CV
机构
*
Guangdong Institute of Intelligence Science and Technology(广东智能科技研究院)
;
Zhejiang University(浙江大学)
;
Southeast University(东南大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Institute of Science Tokyo(东京科学大学)
;
Shanghai Jiao Tong University(上海交通大学)
机构
*
College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院)
;
Department of Bioengineering and Imperial-X, Imperial College London(帝国理工学院伦敦校区生物工程系)
;
Department of Pathology, Xiangtan Maternal and Child Health Hospital(湘潭 maternal and child health hospital pathology department)
;
Department of Pathology, The First People’s Hospital of Xiangtan City(湘潭市第一人民医院病理科)
机构
*
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
Zhejiang Key Laboratory of Digital-Intelligence Service Technology(浙江省数字化服务技术重点实验室)
;
Hong Kong Polytechnic University(香港理工大学)
;
The University of Southern Queensland(南昆士兰大学)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);分类 cs.CV
P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling
P-MTP: 通过渐进深度缩放的多令牌预测实现高效文档解析
Le Xiang, Chenxi Zhai, Shu Wei, Jingjing Wu, Qunyi Xie, Xiao Tan, Kunbin Chen, Wei He
机构
*
Department of Computer Vision Technology (VIS) Baidu Inc China(百度计算机视觉技术部(VIS))
;
Tsinghua University Shenzhen International Graduate School China(清华大学深圳国际研究生院)
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
评估稀疏自编码器与概念标注的可解释性
Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Begüm Demir
机构
*
The Berlin Institute for the Foundations of Learning and Data (BIFOLD)(柏林学习与数据基础研究所)
;
Technische Universität Berlin(柏林工业大学)
;
INRAE(法国国家农业、食品与环境研究院)
;
Inria, EVERGREEN(法国国家信息与自动化研究所,EVERGREEN)
;
UMR TETIS, Univ Montpellier(UMR TETIS,蒙彼利埃大学)
专题命中
VLM训练与架构
:vision language model(abstract);分类 cs.CV、cs.AI