DINORANKCLIP: DINOv3 Distillation and Injection for Vision-Language Pretraining with High-Order Ranking Consistency
DINORANKCLIP:基于DINOv3蒸馏与注入的视觉-语言预训练高阶排名一致性方法
Shuyang Jiang, Nan Yu, Yiming Zhang, Zenghui Ding, Zhenyu Wu
机构
*
University of California, Los Angeles(加州大学洛杉矶分校)
;
Aimaikj(阿米亚克j)
;
HFIPS, Chinese Academy of Sciences(中国科学院HFIPS)
;
University of Science and Technology of China(中国科学技术大学)
;
Chinese Academy of Sciences(中国科学院)
;
National University of Defense Technology(国防科技大学)
机构
*
School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院)
;
School of Computer Science and Technology and the State Key Laboratory of Advanced Rail Autonomous Operation, Beijing Jiaotong University(北京交通大学计算机科学与技术学院和先进轨道交通自主运行国家重点实验室)
;
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
;
School of Mechanical Engineering, Guizhou University(贵州大学机械工程学院)
Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
无人机视角下视觉语言模型能否思考?统一无人机推理与生成
Jintao Sun, Gangyi Ding, Donglin Di, Hu Zhang, Zhedong Zheng
机构
*
Beijing Institute of Technology(北京理工大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
CSIRO Data61(澳大利亚联邦科学与工业研究组织 Data61 分部)
;
University of Macau(澳门大学)
机构
*
School of Computer Science and Technology, Harbin Institute of Technology, China(哈尔滨工业大学计算机科学与技术学院)
;
CFAR and IHPC, Agency for Science, Technology and Research (A*STAR), Singapore(新加坡科技研究局(A*STAR)的CFAR和IHPC)
;
College of Computing and Data Science, Nanyang Technological University, Singapore(新加坡南洋理工大学计算与数据科学学院)