CommentsPeer-reviewed and presented at the 1st Workshop on Toward Trustworthy Vision-Language Models in the Wild (TrustVLM), co-located with ACM ICMR 2026, Amsterdam. Non-archival workshop. Reviews public on OpenReview. 5 pages, 2 figures
机构
*
National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)
;
School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
;
AMAP, Alibaba Group(阿里集团AMAP)
;
School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.LG
Ruiqi Wu, Yuang Yao, Tengfei Ma, Chenran Zhang, Na Su, Tao Zhou, Geng Chen, Wen Fan, Yi Zhou
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Department of Ophthalmology, The First Affiliated Hospital of Nanjing Medical University(南京医科大学第一附属医院眼科学系)
;
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
;
School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
OmniAD:基于多模态推理的工业异常检测与理解
Shifang Zhao, Yiheng Lin, Lu Han, Yao Zhao, Yunchao Wei
机构
*
Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学学院)
;
Visual Intelligence + X International Joint Laboratory of the Ministry of Education(教育部视觉智能+X国际合作实验室)
;
Key Laboratory of Noise and Vibration Research, Institute of Acoustics, Chinese Academy of Sciences(中国科学院声学研究所噪声与振动重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)