Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
用于遥感图像理解的多模态大语言模型:领域特定还是通用?
Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li
机构
*
School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院)
;
Intelligent Game and Decision Lab (IGDL)(智能游戏与决策实验室)
;
Yuelushan Center for Industrial Innovation(岳麓山工业创新中心)
专题命中
视觉问答
:multimodal large language model(title,abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV
D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
D3VL:利用语言模型从3D时间序列数据和视频中理解驾驶场景
Heesang Han, A. Lynn Abbott, Abhijit Sarkar
机构
*
Bradley Department of Electrical and Computer Engineering, Virginia Tech(弗吉尼亚理工大学布拉德利电气与计算机工程系)
;
Virginia Tech Transportation Institute(弗吉尼亚理工大学交通研究所)
;
Sanghani Center for Artificial Intelligence and Data Analytics(桑哈尼人工智能与数据分析中心)
专题命中
视觉问答
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
Pathologist Attention-Aligned Report Generation for Prostate Histopathology
用于前列腺组织病理学的病理学家注意力对齐报告生成
Ruoyu Xue, Suryakant Singh, Souradeep Chakraborty, Pierre Marza, Oksana Yaskiv, Constantin Friedman, Natallia Sheuka, Paul Friedman, Bharat Ramlal, Beatrice Knudsen, Rajarsi Gupta, Joel Saltz, Prateek Prasanna, Gregory Zelinsky, Dimitris Samaras
机构
*
Department of Computer Science, Stony Brook University(纽约州立大学石溪分校计算机科学系)
;
Department of Biomedical Informatics, Stony Brook University(纽约州立大学石溪分校生物医学信息学系)
;
Université Paris-Saclay, CentraleSupélec, Gustave Roussy, INSERM, IHU PRISM, Cancer Data Science Unit(巴黎萨克雷大学、中央理工高等电力学院、古斯塔夫·鲁西研究所、法国国家健康与医学研究院、PRISM综合大学医院、癌症数据科学单元)
;
Université Paris-Saclay, CentraleSupélec, MICS Laboratory(巴黎萨克雷大学、中央理工高等电力学院、MICS实验室)
;
Department of Pathology and Laboratory Medicine, Northwell Health Laboratories(诺斯韦尔健康实验室病理与检验医学部)
;
Department of Pathology, University of Utah School of Medicine(犹他大学医学院病理系)
;
Department of Psychology, Stony Brook University(纽约州立大学石溪分校心理学系)
Comments11 pages, 4 figures, accepted for publication at the 29th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI 2026)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
University of Freiburg(弗莱堡大学)
;
Hangzhou City University(杭州城市大学)
;
XGRIDS(XGRIDS公司)
;
Fudan University(复旦大学)
;
Hong Kong Baptist University(香港浸会大学)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding
KineBench:通过无逆动力学模型的运动学基础对具身世界模型进行基准测试
Zeyu Liu, Zhangzhe Zhu, Yang Zhang, Chenyou Fan, Chenjia Bai, Xuelong Li
机构
*
Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院)
;
National University of Singapore(新加坡国立大学)
;
Fudan University(复旦大学)
;
Tsinghua University(清华大学)
;
Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院)
机构
*
Shandong Second Medical University(山东第二医科大学)
;
Shandong University(山东大学)
;
Bairong Inc.(百融云创科技股份有限公司)
;
Ke Holdings Inc.(贝壳控股有限公司)
;
School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院)
;
Shandong Provincial Key Laboratory of Computing-Network Integration, Shandong University(山东大学计算网络融合技术山东省重点实验室)
专题命中
文档图表理解
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
机构
*
Southeast University(东南大学)
;
Purple Mountain Laboratories(紫金山实验室)
;
Institute of AI for Industries(人工智能产业研究院)
;
Chinese Academy of Sciences(中国科学院)
Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation
基于智能识别与生成的电子元件符号与引脚封装数据库
Yichen Shi, Yuzhi Liu, Zhuofu Tao, Li Huang, Yuhao Gao, Ting-Jung Lin, Lei Hel
机构
*
Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究院,东方理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
BTD Technology(BTD科技)
专题命中
其他VLM
:multimodal large language model(abstract);分类 cs.AI
VizRAG: Enhancing Retrieval-Augmented Generation with Hypergraph Visualization
VizRAG:通过超图可视化增强检索增强生成
Yanbin Wei, Yang Chen, Renling Gan, Ziru Liu, Xinyu Fu, Chun Kang, Ning Lu, Rui Liu, Yu Zhang, James Kwok
机构
*
Southern University of Science and Technology(南方科技大学)
;
Hong Kong University of Science and Technology(香港科技大学)
;
Huawei Research(华为研究院)
;
Beihang University(北京航空航天大学)
专题命中
其他VLM
:multimodal large language model(abstract)