机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
National University of Singapore(新加坡国立大学)
;
Wuhan AI Research(武汉人工智能研究所)
;
Tsinghua University(清华大学)
Semantically Aware UAV Landing Site Assessment from Remote Sensing Imagery via Multimodal Large Language Models
基于多模态大语言模型的语义感知无人机着陆地评估:从遥感图像
Chunliang Hua, Zeyuan Yang, Lei Zhang, Jiayang Sun, Fengwen Chen, Chunlan Zeng, Xiao Hu
机构
*
School of Information Science and Engineering, Southeast University(信息科学与工程学院,东南大学)
;
LASER, International Digital Economy Academy(LASER,国际数字经济学院)
;
Department of Electronic Engineering, East China Normal University(电子工程系,华东师范大学)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI
机构
*
University of Alabama(阿拉巴马大学)
;
University of Maryland, College Park(马里兰大学学院公园分校)
;
Stony Brook University(石溪大学)
;
University of Florida(佛罗里达大学)
;
Johns Hopkins University(约翰霍普金斯大学)
Revisiting KRISP: A Lightweight Reproduction and Analysis of Knowledge-Enhanced Vision-Language Models
重新审视 KRISP:一种轻量级的知识增强视觉-语言模型的再现与分析
Souradeep Dutta, Keshav Bulia, Neena S Nair
机构
*
Centre For Digital Health(数字健康中心)
;
Indian Institute of Technology Bombay(孟买印度理工学院)
;
Department of Metallurgical engineering & Material science(冶金工程与材料科学系)
;
Department of Bioscience & Bioengineering(生物科学与生物工程系)
GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model
GeoReasoner: 基于大规模视觉-语言模型的街景地理定位
Ling Li, Yu Ye, Yao Zhou, Bingchuan Jiang, Wei Zeng
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Tongji University(同济大学)
;
Independent Researcher(独立研究者)
;
Information Engineering University(信息工程大学)