Improving Vision-language Models with Perception-centric Process Reward Models
通过以感知为中心的过程奖励模型改进视觉-语言模型
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) ; Bytedance(字节跳动) ; University of California, San Diego(加州大学圣地亚哥分校) ; The Hong Kong University of Science and Technology(香港科技大学)
AI总结 本文提出Perceval模型,通过token级错误定位提升视觉-语言模型的推理能力,通过感知驱动的监督策略实现细粒度训练与推理优化,实验显示在多个领域基准上显著提升性能。
Comments 8 pages
Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 33099-33109