CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference
CRISP:用于高效LVLM推理的预LLM文本驱动视觉令牌剪枝
机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) ; School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院) ; School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院) ; College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)
专题命中 效率与部署 :LLM(title,title_cn);language model(abstract)
AI总结 针对大型视觉语言模型推理开销大的问题,提出CRISP框架,通过文本驱动在预LLM阶段剪枝视觉令牌,分两阶段工作,实验表明其在激进剪枝率下能保持高性能,降低推理成本和延迟,是高效LVLM推理的实用方案。
Comments Accepted by the 2026 IEEE International Conference on Multimedia and Expo (ICME 2026) as Oral