UniPCB: A Unified Vision-Language Benchmark for Open-Ended PCB Quality Inspection
UniPCB: 一个统一的视觉-语言基准用于开放式PCB质量检测
Fuxiang Sun, Xi Jiang, Jiansheng Wu, Haigang Zhang, Feng Zheng, Jinfeng Yang
机构
*
Shenzhen Polytechnic University(深圳职业技术大学)
;
Southern University of Science and Technology(南方科技大学)
;
University of Science and Technology Liaoning(辽宁科技大学)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构
*
Ministry of Education Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(教育部图像处理与智能控制重点实验室,人工智能与自动化学院,华中科技大学)
;
State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences(脑认知与脑启发智能技术国家重点实验室,自动化研究所,中国科学院)
Measuring the Unspoken: A Disentanglement Model and Benchmark for Psychological Analysis in the Wild
测量未言之: 一种解耦模型和基准用于野外心理分析
Yigui Feng, Qinglin Wang, Haotian Mo, Yang Liu, Ke Liu, Gencheng Liu, Xinhai Chen, Siqi Shen, Songzhu Mei, Jie Liu
机构
*
College of Computer Science, National University of Defense Technology(计算机科学学院,国防科技大学)
;
Shien-Ming Wu School of Intelligent Engineering, South China University of Technology(智能工程学院,华南理工大学)
Does Bigger Mean Better? Comparitive Analysis of CNNs and Biomedical Vision Language Modles in Medical Diagnosis
Ran Tong, Jiaqi Liu, Tong Wang, Xin Hu, Su Liu, Lanruo Wang, Jiexi Xu
机构
*
University of Texas at Dallas(德克萨斯大学达拉斯分校)
;
Independent Researcher(独立研究者)
;
Duke University(杜克大学)
;
University of Michigan Ann Arbor(密歇根大学安娜堡分校)
;
Georgia Institute of Technology(佐治亚理工学院)
;
University of California, Irvine(加州大学 Irvine 分校)
Laugh, Relate, Engage: Stylized Comment Generation for Short Videos
Xuan Ouyang, Senan Wang, Bouzhou Wang, Siyuan Xiahou, Jinrong Zhou, Yuekang Li
机构
*
University of New South Wales(新南威尔士大学)
;
University of Sydney(悉尼大学)
;
The University of Hong Kong(香港大学)
;
University of Southern California(南加州大学)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG
Generating Accurate and Detailed Captions for High-Resolution Images
Hankyeol Lee, Gawon Seo, Kyounggyu Lee, Dogun Kim, Kyungwoo Song, Jiyoung Jung
机构
*
Department of Artificial Intelligence, University of Seoul(首尔大学人工智能系)
;
Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系)
;
Department of Applied Statistics, Yonsei University(延世大学应用统计系)
机构
*
Zhejiang Gongshang University(浙江工商大学)
;
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室)
;
Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波)
;
Meituan Inc.(美团公司)
;
National University of Singapore(新加坡国立大学)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI