AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
Shixiong Xu, Chenghao Zhang, Lubin Fan, Yuan Zhou, Bin Fan, Shiming Xiang, Gaofeng Meng, Jieping Ye
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室)
;
CASIA(中国科学院自动化所)
;
Alibaba Cloud(阿里云)
;
School of Intelligence Science and Technology(智能科学与技术学院)
;
CAIR(中国科学院香港创新研究院)