SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
SCAM:多模态基础模型的现实世界字形鲁棒性评估
Justus Westerhoff, Erblina Purelku, Jakob Hackstein, Jonas Loos, Leo Pinetzki, Erik Rodner, Lorenz Hufe
机构
*
BLISS e.V.(BLISS协会)
;
Berliner Hochschule für Technik (BHT)(柏林技术大学)
;
Technische Universität Berlin(柏林技术大学)
;
KI Werkstatt, Hochschule für Technik und Wirtschaft Berlin (HTW)(柏林技术经济学院)
;
Merantix Momentum(Merantix Momentum公司)
;
Fraunhofer Heinrich-Hertz-Institut, Berlin, Germany(柏林弗劳恩霍夫 Heinrich-Hertz 研究所)
专题命中
多模态评测
:multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV、cs.AI
MLlm-DR: Towards Explainable Depression Recognition with MultiModal Large Language Models
MLlm-DR: 向多模态大语言模型的可解释性抑郁识别迈进
Wei Zhang, Juan Chen, En Zhu, Wenhong Cheng, YunPeng Li, Yanbo J. Wang
机构
*
National University of Defense Technology(国防科技大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Shanghai Mental Health Center, Shanghai Jiao Tong University School of Medicine(上海精神卫生中心,上海交通大学医学院)
;
Nanjing Industria Tenebris Information Technology Co., Ltd.(南京Industria Tenebris信息技术有限公司)
;
National University of Uzbekistan named after Mirzo Ulugbek(乌兹别克斯坦Mirzo Ulugbek命名的国立大学)
UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
UniSAFE:统一多模态模型安全评估的综合基准
Segyu Lee, Boryeong Cho, Hojung Jung, Seokhyun An, Juhyeong Kim, Jaehyun Kwak, Yongjin Yang, Sangwon Jang, Youngrok Park, Wonjun Chang, Se-Young Yun
机构
*
KAIST AI(KAIST人工智能研究院)
;
Department of Computer Science and Engineering, UNIST(UNIST计算机科学与工程系)
;
Department of Mathematical Sciences, KAIST(KAIST数学科学系)
;
University of Toronto(多伦多大学)
;
KAIST CS(KAIST计算机科学系)
Comments5 tables, 6 figures, Submitted to International Conference on Power, Electronics, Communications, Computing, and Intelligent Infrastructure 2026
LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
LMOD+: 一个全面的多模态数据集和基准,用于开发和评估眼科中的多模态大语言模型
Zhenyue Qin, Yang Liu, Yu Yin, Jinyu Ding, Haoran Zhang, Anran Li, Dylan Campbell, Xuansheng Wu, Ke Zou, Tiarnan D. L. Keenan, Emily Y. Chew, Zhiyong Lu, Yih Chung Tham, Ninghao Liu, Xiuzhen Zhang, Qingyu Chen
机构
*
School of Medicine, Yale University(耶鲁大学医学院)
;
School of Computing, Australian National University(澳大利亚国立大学计算机学院)
;
School of Engineering, Imperial College London(伦敦帝国理工学院工程学院)
;
School of Computing, University of Georgia(佐治亚大学计算机学院)
;
Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学杨秀隆医学学院)
;
National Eye Institute, National Institutes of Health(美国国立卫生研究院眼科研究所)
;
National Library of Medicine, National Institutes of Health(美国国立卫生研究院国家医学图书馆)
;
School of Computing Technologies, RMIT University(皇家墨尔本理工大学计算机技术学院)
Concept-to-Pixel: Prompt-Free Universal Medical Image Segmentation
概念到像素:无提示通用医学图像分割
Haoyun Chen, Fenghe Tang, Wenxin Ma, Shaohua Kevin Zhou
机构
*
School of Biomedical Engineering, Division of Life Sciences
;
Medicine, University of Science
;
Technology of China (USTC), Hefei, Anhui 230026, China Center for Medical Imaging, Robotics, Analytic Computing \& Learning (MIRACLE), Suzhou Institute for Advanced Research, USTC, Suzhou, Jiangsu 215123, China Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology, Suzhou Jiangsu, 215123, China State Key Laboratory of Precision
A Comprehensive Benchmark of Histopathology Foundation Models for Kidney Digital Pathology Images
一种针对肾数字病理图像的病理基础模型综合基准测试
Harishwar Reddy Kasireddy, Patricio S. La Rosa, Akshita Gupta, Anindya S. Paul, Jamie L. Fermin, William L. Clapp, Meryl A. Waldman, Tarek M. El-Ashkar, Sanjay Jain, Luis Rodrigues, Kuang Yu Jen, Avi Z. Rosenberg, Michael T. Eadon, Jeffrey B. Hodgin, Pinaki Sarder
机构
*
Department of Electrical and Computer Engineering, University of Florida(佛罗里达大学电气与计算机工程系)
;
Seed Production Innovation, Crop Science Division, Bayer Company(拜耳公司种子生产创新部)
;
Division of Medicine – Quantitative Health, University of Florida(佛罗里达大学医学部-定量健康分部)
;
Department of Health Outcomes and Biomedical Informatics, University of Florida(佛罗里达大学健康结果与生物医学信息学系)
;
Department of Pathology, Immunology and Laboratory Medicine, University of Florida College of Medicine(佛罗里达大学医学院病理学、免疫学与实验室医学系)
;
Kidney Disease Branch, National Institute of Diabetes and Digestive and Kidney Diseases, National Institutes of Health(美国国立卫生研究院糖尿病、消化系统与肾病研究所肾病分支)
;
Indiana University School of Medicine(印第安纳大学医学院)
;
Departments of Medicine, Washington University School of Medicine(华盛顿大学医学院医学部)
;
Universidade de Coimbra(科英布拉大学)
;
Department of Pathology and Laboratory Medicine, University of California at Davis School of Medicine(加州大学戴维斯分校医学院病理学与实验室医学系)
;
Department of Pathology, Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院病理学系)
专题命中
多模态评测
:multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
FINER:MLLMs在细粒度负查询下产生幻觉
Rui Xiao, Sanghwan Kim, Yongqin Xian, Zeynep Akata, Stephan Alaniz
机构
*
Technical University of Munich(慕尼黑技术大学)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
Helmholtz Munich(海德堡-慕尼黑亥姆霍尔茨中心)
;
Google(谷歌)
;
LTCI, Télécom Paris, Institut Polytechnique de Paris(LTCI,巴黎电信学院,巴黎理工学院)
VL-RouterBench: A Benchmark for Vision-Language Model Routing
VL-RouterBench:一种用于视觉-语言模型路由的基准测试
Zhehao Huang, Baijiong Lin, Jingyuan Zhang, Jingying Wang, Yuhang Liu, Ning Lu, Tao Li, Xiaolin Huang
机构
*
Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University(1 图像处理与模式识别研究所,上海交通大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(2 香港科学与技术大学(广州))
;
The Hong Kong University of Science and Technology(3 香港科学与技术大学)
UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detection
UAV-CB: 一种复杂背景RGB-T数据集和局部频率桥梁网络用于UAV检测
Shenghui Huang, Menghao Hu, Longkun Zou, Hongyu Chi, Zekai Li, Feng Gao, Fan Yang, Qingyao Wu, Ke Chen
机构
*
Pengcheng Laboratory(鹏城实验室)
;
South China University of Technology(南方科技大学)
;
Peking University(北京大学)
;
Xinjiang University(新疆大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
DexViTac: Collecting Human Visuo-Tactile-Kinematic Demonstrations for Contact-Rich Dexterous Manipulation
DexViTac:收集人类视觉-触觉-运动示范以实现富接触的灵巧操作
Xitong Chen, Yifeng Pan, Min Li, Xiaotian Ding
机构
*
State Key Laboratory of Intelligent Manufacturing Equipment and Technology(智能制造装备与技术国家重点实验室)
;
Huazhong University of Science and Technology(华中科技大学)
;
Wuhan Huaweike Intelligent Technology Co., Ltd.(武汉华为凯科技有限公司)