Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
用于遥感图像理解的多模态大语言模型:领域特定还是通用?
Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li
机构
*
School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院)
;
Intelligent Game and Decision Lab (IGDL)(智能游戏与决策实验室)
;
Yuelushan Center for Industrial Innovation(岳麓山工业创新中心)
专题命中
视觉问答
:multimodal large language model(title,abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV
D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
D3VL:利用语言模型从3D时间序列数据和视频中理解驾驶场景
Heesang Han, A. Lynn Abbott, Abhijit Sarkar
机构
*
Bradley Department of Electrical and Computer Engineering, Virginia Tech(弗吉尼亚理工大学布拉德利电气与计算机工程系)
;
Virginia Tech Transportation Institute(弗吉尼亚理工大学交通研究所)
;
Sanghani Center for Artificial Intelligence and Data Analytics(桑哈尼人工智能与数据分析中心)
专题命中
视觉问答
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
Pathologist Attention-Aligned Report Generation for Prostate Histopathology
用于前列腺组织病理学的病理学家注意力对齐报告生成
Ruoyu Xue, Suryakant Singh, Souradeep Chakraborty, Pierre Marza, Oksana Yaskiv, Constantin Friedman, Natallia Sheuka, Paul Friedman, Bharat Ramlal, Beatrice Knudsen, Rajarsi Gupta, Joel Saltz, Prateek Prasanna, Gregory Zelinsky, Dimitris Samaras
机构
*
Department of Computer Science, Stony Brook University(纽约州立大学石溪分校计算机科学系)
;
Department of Biomedical Informatics, Stony Brook University(纽约州立大学石溪分校生物医学信息学系)
;
Université Paris-Saclay, CentraleSupélec, Gustave Roussy, INSERM, IHU PRISM, Cancer Data Science Unit(巴黎萨克雷大学、中央理工高等电力学院、古斯塔夫·鲁西研究所、法国国家健康与医学研究院、PRISM综合大学医院、癌症数据科学单元)
;
Université Paris-Saclay, CentraleSupélec, MICS Laboratory(巴黎萨克雷大学、中央理工高等电力学院、MICS实验室)
;
Department of Pathology and Laboratory Medicine, Northwell Health Laboratories(诺斯韦尔健康实验室病理与检验医学部)
;
Department of Pathology, University of Utah School of Medicine(犹他大学医学院病理系)
;
Department of Psychology, Stony Brook University(纽约州立大学石溪分校心理学系)
Comments11 pages, 4 figures, accepted for publication at the 29th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI 2026)
机构
*
School of Mechanical Engineering, Beijing Institute of Technology(北京理工大学机械工程学院)
;
National Engineering Research Center of Electric Vehicles, Beijing Institute of Technology(北京理工大学电动车辆国家工程研究中心)
;
School of Mechanical Engineering, Southeast University(东南大学机械工程学院)
机构
*
Hunan University, Changsha, China(湖南大学)
;
Wuhan University of Technology, Wuhan, China(武汉理工大学)
;
Huazhong University of Science and Technology, Wuhan, China(华中科技大学)
专题命中
视觉推理
:VLM(abstract,abstract_cn);vision language model(abstract);分类 cs.AI
Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning
不要愚弄我两次:通过经验驱动推理在野外适应逆境
Navin Sriram Ravie, Andrew Jong, Krrish Jain, John Liu, Omar Alama, Bijo Sebastian, Sebastian Scherer
机构
*
Department of Engineering Design, Indian Institute of Technology, Madras(印度理工学院工程设计系,马德拉斯)
;
Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
University of Freiburg(弗莱堡大学)
;
Hangzhou City University(杭州城市大学)
;
XGRIDS(XGRIDS公司)
;
Fudan University(复旦大学)
;
Hong Kong Baptist University(香港浸会大学)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中
视觉定位与Grounding
:MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.AI
KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding
KineBench:通过无逆动力学模型的运动学基础对具身世界模型进行基准测试
Zeyu Liu, Zhangzhe Zhu, Yang Zhang, Chenyou Fan, Chenjia Bai, Xuelong Li
机构
*
Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院)
;
National University of Singapore(新加坡国立大学)
;
Fudan University(复旦大学)
;
Tsinghua University(清华大学)
;
Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院)
机构
*
School of Information Science and Technology, Yunnan Normal University(云南师范大学信息科学与技术学院)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
;
College of Computer Science, Beijing University of Technology(北京工业大学计算机学院)
机构
*
Shandong Second Medical University(山东第二医科大学)
;
Shandong University(山东大学)
;
Bairong Inc.(百融云创科技股份有限公司)
;
Ke Holdings Inc.(贝壳控股有限公司)
;
School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院)
;
Shandong Provincial Key Laboratory of Computing-Network Integration, Shandong University(山东大学计算网络融合技术山东省重点实验室)
专题命中
文档图表理解
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV