3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects
3D-缺陷基准:用于细粒度3D生成缺陷的视觉语言模型评估管道的控制因子研究
Zhenyu Zhao, Nanshan Jia, Jihyeon Je, Yifu Tang, Alvin Chan, Michael Spedden, Michael V. Palleschi, Sui Huang, Jingshen Wang, Zeyu Zheng
机构
*
Roblox Corporation(罗布乐思公司)
;
Berkeley AI Research Lab & Department of Industrial Engineering and Operations Research, University of California, Berkeley(加州大学伯克利分校伯克利人工智能研究实验室及工业工程与运筹学系)
;
Computer Science Department, Stanford University(斯坦福大学计算机科学系)
;
Division of Biostatistics, University of California, Berkeley(加州大学伯克利分校生物统计学系)
MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models
MMBU: 大规模多模态生物医学理解基准,用于探测视觉语言模型的感知能力
Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin, Min Woo Sun, Josiah Aklilu, James Burgess, Yuhui Zhang, Ryan Nayebi, Paola Avila, Robayo, Jin Ye, Ming Hu, Zhongying Deng, Junjun He, Xin Chen, Yue Yao, Robert Tibshirani, Jeffrey J. Nirschl, Serena Yeung-Levy
机构
*
Stanford University(斯坦福大学)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Instituto Tecnológico de Monterrey(蒙特雷技术学院)
;
Monash University(墨尔本大学)
;
University of Cambridge(剑桥大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Shandong University(山东大学)
Vision-Language Models as Zero-Annotation Oracles in Histopathology
视觉-语言模型作为组织病理学中的零标注预言机
Vishal Jain, Giorgio Buzzanca, Sarah Cechnicka, Maarten Naesens, Priyanka Koshy, Tri Nguyen, Jesper Kers, Candice Roufosse, Bernhard Kainz
机构
*
Imperial College London(帝国理工学院)
;
Leiden University Medical Center(莱顿大学医学中心)
;
KU Leuven(鲁汶大学)
;
University Hospitals Leuven(鲁汶大学医院)
;
University Medical Center Utrecht(乌得勒支大学医学中心)
;
Friedrich-Alexander University Erlangen-Nürnberg(埃尔朗根-纽伦堡大学)
机构
*
School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院,中国)
;
Alibaba Group(阿里巴巴集团)
;
School of Computer Science and Engineering, School of Intelligence Science and Engineering, and Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Southeast University, China(东南大学计算机科学与工程学院、智能科学与工程学院以及新一代人工智能技术及其交叉应用关键实验室,中国)
;
Wangxuan Institute of Computer Technology, National Key Laboratory for Multimedia Information Processing, Peking University, China(北京大学王轩计算机技术研究所、多媒体信息处理国家重点实验室,中国)
;
University of Copenhagen, Denmark(丹麦哥本哈根大学)
Autonomous VR-Based Risk Detection for Situational Awareness in Dangerous Settings
基于自主虚拟现实的危险环境态势感知风险检测
Mohammad Eskandari, Murali Krishna Varma Indukuri, Stephanie M. Lukin, Cynthia Matuszek
机构
*
Interactive Robotics and Language Lab, University of Maryland Baltimore County(马里兰大学巴尔的摩县分校交互式机器人与语言实验室)
;
DEVCOM Army Research Laboratory(陆军研究实验室)
专题命中
视觉定位与Grounding
:VLM(summary_cn,abstract);vision language model(abstract);分类 cs.CV、cs.AI
Vision-language models for chest radiography do not always need the image
胸部X光片的视觉-语言模型并不总是需要图像
Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
机构
*
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃尔朗根-纽伦堡大学模式识别实验室)
;
Department of Diagnostic and Interventional Radiology, TUM University Clinic, School of Medicine and Health, Klinikum rechts der Isar, Technical University of Munich(慕尼黑工业大学医学院与健康学院伊萨尔河右岸医院诊断与介入放射学系)
;
Lab for AI in Medicine, RWTH Aachen University(亚琛工业大学医学人工智能实验室)
;
Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen(亚琛工业大学医院诊断与介入放射学系)