VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
VLA-InfoEntropy:一种无需训练的视觉-注意力信息熵方法,用于视觉-语言-动作模型的推理加速与成功
机构 * Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司) ; Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);action model(title);分类 cs.RO、cs.CV
AI总结 本文提出VLA-InfoEntropy方法,通过图像熵和注意力熵结合时间步信息,动态调整模型关注区域,减少冗余并提升推理效率。
Comments Accepted to the 2026 IEEE International Conference on Multimedia and Expo (ICME 2026)