Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding
面向低延迟视觉语言模型:在自我中心视觉理解中实现双重正确预测
机构 * Dolby Laboratories, Inc.(杜比实验室公司) ; University of Delaware(特拉华大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV
AI总结 针对人机协作中低延迟需求,提出基于双重正确预测的剪枝策略,在保持证据定位的同时提升预测准确性,在自我中心视频数据集上实现最高精度与双重正确性。
Comments International Conference on Intelligent Robots and Systems (IROS) 2026