Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
用于遥感图像理解的多模态大语言模型:领域特定还是通用?
Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li
机构
*
School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院)
;
Intelligent Game and Decision Lab (IGDL)(智能游戏与决策实验室)
;
Yuelushan Center for Industrial Innovation(岳麓山工业创新中心)
专题命中
视觉问答
:multimodal large language model(title,abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV
D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
D3VL:利用语言模型从3D时间序列数据和视频中理解驾驶场景
Heesang Han, A. Lynn Abbott, Abhijit Sarkar
机构
*
Bradley Department of Electrical and Computer Engineering, Virginia Tech(弗吉尼亚理工大学布拉德利电气与计算机工程系)
;
Virginia Tech Transportation Institute(弗吉尼亚理工大学交通研究所)
;
Sanghani Center for Artificial Intelligence and Data Analytics(桑哈尼人工智能与数据分析中心)
专题命中
视觉问答
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
Pathologist Attention-Aligned Report Generation for Prostate Histopathology
用于前列腺组织病理学的病理学家注意力对齐报告生成
Ruoyu Xue, Suryakant Singh, Souradeep Chakraborty, Pierre Marza, Oksana Yaskiv, Constantin Friedman, Natallia Sheuka, Paul Friedman, Bharat Ramlal, Beatrice Knudsen, Rajarsi Gupta, Joel Saltz, Prateek Prasanna, Gregory Zelinsky, Dimitris Samaras
机构
*
Department of Computer Science, Stony Brook University(纽约州立大学石溪分校计算机科学系)
;
Department of Biomedical Informatics, Stony Brook University(纽约州立大学石溪分校生物医学信息学系)
;
Université Paris-Saclay, CentraleSupélec, Gustave Roussy, INSERM, IHU PRISM, Cancer Data Science Unit(巴黎萨克雷大学、中央理工高等电力学院、古斯塔夫·鲁西研究所、法国国家健康与医学研究院、PRISM综合大学医院、癌症数据科学单元)
;
Université Paris-Saclay, CentraleSupélec, MICS Laboratory(巴黎萨克雷大学、中央理工高等电力学院、MICS实验室)
;
Department of Pathology and Laboratory Medicine, Northwell Health Laboratories(诺斯韦尔健康实验室病理与检验医学部)
;
Department of Pathology, University of Utah School of Medicine(犹他大学医学院病理系)
;
Department of Psychology, Stony Brook University(纽约州立大学石溪分校心理学系)
Comments11 pages, 4 figures, accepted for publication at the 29th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI 2026)