Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
机构 * Tel Aviv University(特拉维夫大学) ; Google Research(谷歌研究)
Comments Accepted to NAACL 2025
Journal ref Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics, Human Language Technologies, Long Papers, pp. 10497-10518