Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
CommentsSubmitted to the interactivity track of the 21st ACM/IEEE International Conference on Human-Robot Interaction on December 2025, accepted January 2026
Journal refHRI Companion 2026: Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction
机构
*
Department of Computer Science, Purdue University(普渡大学计算机科学系)
;
Department of Computer Science, Rice University(莱斯大学计算机科学系)
;
Ken Kennedy Institute, Rice University(莱斯大学肯尼迪研究所)
MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine
MedGPT-oss: 为生物医学训练一个通用的视觉-语言模型
Kai Zhang, Zhengqing Yuan, Cheng Peng, Songlin Zhao, Mengxian Lyu, Ziyi Chen, Yanfang Ye, Wei Liu, Ying Zhang, Kaleb E Smith, Lifang He, Lichao Sun, Yonghui Wu
机构
*
Department of Computer Science and Engineering, Lehigh University(莱斯大学计算机科学与工程系)
;
Department of Computer Science and Engineering, University of Notre Dame(圣母大学计算机科学与工程系)
;
Department of Health Outcomes & Biomedical Informatics, University of Florida(佛罗里达大学健康结果与生物医学信息学系)
;
Department of Radiation Oncology, Mayo Clinic(梅奥诊所放射肿瘤科)
;
Research Computing, University of Florida(佛罗里达大学研究计算中心)
;
AI Technology Center, NVIDIA(NVIDIA人工智能技术中心)