HPE-CogVLM: Advancing Vision Language Models with a Head Pose Grounding Task
HPE-CogVLM: 通过头部姿态接地任务提升视觉语言模型
机构 * Docomo Innovations, Inc.(Docomo创新公司) ; Department of Computer Science and Engineering, Santa Clara University(圣克拉拉大学计算机科学与工程系)
专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.CV、cs.AI
AI总结 本文提出HPE-CogVLM框架,利用VLM的物体检测能力提升头部姿态估计精度,通过改进的LoRA层合并方法,有效解决融合任务中的响应格式问题,实现优于现有方法的性能。
Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology (TCSVT), 2026. This version includes major updates in methodology and experiments. The final version is available at IEEE Xplore
Journal ref IEEE Transactions on Circuits and Systems for Video Technology, Early Access, 2026