Evaluation of Medical Vision Language Models HuluMed and MedGemma, and general purpose chatbots Gemma 3, ChatGPT Plus, and Claude Pro on real previously unseen wound images
医学视觉语言模型 HuluMed 和 MedGemma 以及通用聊天机器人 Gemma 3、ChatGPT Plus 和 Claude Pro 在真实未见伤口图像上的评估
Yunzhe Xue, Mohammed Saim Ahmed Quadri, Neal Panse, Justin W. Ady, Usman Roshan
机构
*
Department of Computer Science, New Jersey Institute of Technology(新泽西理工学院计算机科学系)
;
Vascular and Endovascular Surgery, Robert Wood Johnson Hospital(罗伯特·伍德·约翰逊医院血管外科)
;
Department of Data Science, New Jersey Institute of Technology(新泽西理工学院数据科学系)
专题命中
视觉推理
:vision language model(title);vision-language model(abstract);VLM(abstract_cn);分类 cs.CV
AI总结
本研究评估了六种视觉语言模型在慢性伤口分析任务上的表现,发现通用模型 ChatGPT 和 Claude 显著优于医学专用模型,表明广泛的多模态推理能力比领域知识更重要。
HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models
HANCLIP:双曲角否定视觉语言模型系列
Hoang-Bao Le, Aiden Durrant, Thai Son Mai, Binh T. Nguyen, Liting Zhou, Cathal Gurrin
机构
*
ADAPT Centre Dublin City University, Ireland(爱尔兰都柏林城市大学ADAPT中心)
;
University of East Anglia Norwich, UK(英国东英吉利大学)
;
Queen’s University Belfast Belfast, UK(英国贝尔法斯特女王大学)
;
University of Science Vietnam National University Ho Chi Minh City, Vietnam(越南胡志明市国家大学理科大学)
专题命中
视觉推理
:vision language model(title);vision-language model(abstract);VLM(abstract_cn);分类 cs.CV
TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs
TriViewBench: 多视图结构推理中受控复杂度缩放
Yu-Yang Chen, Lan-Zhe Guo
机构
*
School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
;
National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
专题命中
视觉推理
:visual reasoning(abstract);visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract_cn)
OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models
OpenMedReason: 医学视觉语言模型的科学推理监督
Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
机构
*
York University(约克大学)
;
Vector Institute(向量研究所)
;
University of British Columbia(不列颠哥伦比亚大学)
;
University of Toronto(多伦多大学)
;
Unity Health Toronto / St. Michael’s Hospital(多伦多联合健康/圣迈克尔医院)
;
University Health Network(大学健康网络)
;
Arc Institute(弧研究所)
;
Queen's University(女王大学)
Journal refS. Oh, J. U. Kim and S. Lee, "Rationale-Guided Learning for Multimodal Emotion Recognition," ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2026, pp. 12577-12581
Comments7 pages, 5 figures, 6 tables. Accepted to the 14th IEEE International Conference on Intelligent Mobile Computing (IEEE IMC 2026), Fukuoka, Japan, July 27-30, 2026
PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology
PanDent:面向牙科放射学中全面的牙级结构-语言一致性
Xiaohan Li, Xinyu Liu, Chang Liu, Sum Wing Au Yeung, Jun Liu, Yixuan Yuan, Hui Chen
机构
*
Faculty of Dentistry, The University of Hong Kong(香港大学牙医学院)
;
Imperial College London(帝国理工学院)
;
University of Science and Technology of China(中国科学技术大学)
;
Department of Data and Systems Engineering, The University of Hong Kong(香港大学数据与系统工程系)
;
Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系)
专题命中
视觉推理
:MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.CV
ERQA-Plus: A Diagnostic Benchmark for Reasoning in Embodied AI
ERQA-Plus:具身AI推理的诊断基准
Hong Yang, Basura Fernando
机构
*
Centre for Frontier AI Research, Agency for Science, Technology and Research(新加坡科技研究局前沿人工智能研究中心)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)