Object Tokens as a Bridge Between Segmentation and Visual Question Answering in Robotic Surgery
对象标记作为机器人手术中分割与视觉问答的桥梁
Yiping Li, Ronald de Jong, Romy van Jaarsveld, Franco Badaloni, Gino Kuiper, Jelle Ruurda, Josien Pluim, Marcel Breeuwer
机构
*
Department of Biomedical Engineering, Eindhoven University of Technology(埃因霍温理工大学生物医学工程系)
;
Department of Electrical Engineering, Eindhoven University of Technology(埃因霍温理工大学电气工程系)
;
Department of Surgery, University Medical Center Utrecht(乌得勒支大学医学中心外科)
A Benchmark for Hallucination Detection in VLMs for Gastrointestinal Endoscopy
胃肠内窥镜中视觉语言模型幻觉检测的基准测试
Aminu Lawal, Niyoj Oli, Sachin Acharya, Prashnna Gyawali, Maria Carmen Romano, Binod Bhattarai
机构
*
University of Aberdeen(阿伯丁大学)
;
Nepal Applied Mathematics and Informatics Institute for Research(尼泊尔应用数学与信息学研究所)
;
West Virginia University(西弗吉尼亚大学)
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
波兰医学视觉问答:视觉语言模型未充分利用视觉证据
Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa
机构
*
ARAAI Poland(波兰ARAAI)
;
NASK National Research Institute(NASK国家研究院)
;
Poznań University of Medical Sciences(波兹南医科大学)
;
T. Marciniak Lower Silesian Specialist Hospital(T.马尔奇尼亚克下西里西亚专科医院)
;
Adam Mickiewicz University(亚当·密茨凯维奇大学)
Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models
位置,而非来源:在医学视觉-语言模型中分离推理中介与逢迎行为
Supratik Bhowal, Subhrajyoti Basu, Aritra Gir Mahanta, Anik Pal Chowdhury
机构
*
IEM Kolkata(印度工程管理学院 Kolkata校区)
;
School of UEMK Kolkata(UEMK Kolkata学院)
;
Heritage Institute of Technology Kolkata(加尔各答遗产技术学院)
;
Indian Institute of Information Technology, Kalyani(卡利亚尼印度信息技术学院)
Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks
Earth-OneVision:将遥感多模态大语言模型扩展到更多传感器模态和任务
Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou, Yin Zhuang, Tong Zhang, Hao Wang, He Chen, Jun Li
机构
*
National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing (SBIIP), Beijing Institute of Technology(北京理工大学空间智能信息处理国家重点实验室)
;
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院)
;
Key Laboratory of Technology in Geo-Spatial Information Processing and Application System, Chinese Academy of Sciences(中国科学院地理空间信息处理与应用系统技术重点实验室)
;
Advanced Research Institute of Multidisciplinary Sciences, Beijing Institute of Technology(北京理工大学前沿交叉科学研究院)
;
School of Mechatronical Engineering, Beijing Institute of Technology(北京理工大学机电学院)
;
School of Earth and Space Sciences, Peking University(北京大学地球与空间科学学院)
;
School of Electronics, Peking University(北京大学电子学院)
;
School of Computer Science and Hubei Key Laboratory of Intelligent Geo-Information Processing(华中科技大学计算机科学与技术学院&湖北省智能地理信息处理重点实验室)
专题命中
视觉问答
:MLLM(summary_cn,abstract);multimodal large language model(title);grounding(abstract);分类 cs.CV、cs.AI
机构
*
National and Local Joint Engineering Laboratory of Computer Aided Design, School of Software Engineering, Dalian University(大连大学软件工程学院计算机辅助设计国家地方联合工程实验室)
;
School of Computer Science and Technology, Dalian University of Technology(大连理工大学计算机科学与技术学院)
Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology
面向放射学的空间定位2D视觉-语言模型的可扩展训练
Yusuf Salcan, Simon Ging, Robin Tibor Schirrmeister, Philipp Arnold, Elmar Kotter, Behzad Bozorgtabar, Thomas Brox
机构
*
Computer Vision Group, University of Freiburg, Germany(德国弗莱堡大学计算机视觉组)
;
Department of Radiology, Medical Center -- University of Freiburg, Germany(德国弗莱堡大学医学中心放射科)
;
CRIION-AI Lab, Freiburg, Germany(德国弗莱堡CRIION-AI实验室)
DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery
DM-KG:一种提升街景图像中视觉语言模型空间认知的新方法
Xinyue Xu, Zheng Zhang, Kunyang Ma, Ge Zhu, Lianshuai Cao, Lei Wang, Zixuan Li, Yi Cheng
机构
*
Institute of Surveying and Mapping, Information Engineering University(信息工程大学测绘学院)
;
Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences(中国科学院地理科学与资源研究所)
Search-based Testing of Vision Language Models for In-Car Scene Understanding
基于搜索的车内场景理解视觉语言模型测试
Lev Sorokin, Chen Yang, Ken E. Friedl, Andrea Stocco
机构
*
BMW Group, Technical University of Munich(宝马集团、慕尼黑技术大学)
;
Technical University of Munich(慕尼黑技术大学)
;
Technical University of Munich, fortiss GmbH(慕尼黑技术大学、fortiss GmbH)
专题命中
视觉问答
:VLM(summary_cn,abstract_cn);vision language model(title);vision-language model(abstract);分类 cs.CV