Attention-Steered Vision-Language Models for Sign Language Translation
用于手语翻译的注意力引导视觉-语言模型
机构 * Rochester Institute of Technology(罗切斯特理工学院) ; University of Virginia(弗吉尼亚大学) ; National Technical Institute for the Deaf(国家聋人技术学院)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV
AI总结 针对视觉-语言模型在手语翻译中时空视觉定位差的问题,提出AttnSign框架,通过空间注意力监督和RL运动节奏引导提升性能,在How2Sign和OpenASL基准上表现优于现有方法。