Multi-Modal Semantic Segmentation of Electrolyzer Components for Sustainable Hydrogen Technologies: A Dual-Branch Deep Learning Approach
用于可持续氢能技术的电解槽组件多模态语义分割:一种双分支深度学习方法
Wasimul Karim, Nur Mohammad Fahad, Abdul Hasib Siddique, Md Rafiqul Islam, Hooman Mehdizadeh-Rad, Asif Karim, Sami Azam
机构
*
Applied Artificial Intelligence and INtelligent Systems (AAIINS) Laboratory(应用人工智能与智能系统(AAIINS)实验室)
;
University of Scholars(学者大学)
;
Murdoch University(莫道克大学)
;
Charles Darwin University(查尔斯达尔文大学)
Paired Uterine Whole-Slide Images and Pathology Reports for Multimodal Computational Pathology
用于多模态计算病理学的配对子宫全切片图像和病理报告
Han Li, Jingsong Liu, Ayako Ura, Junlin Hou, Zhengyang Xu, Azar Kazemi, Oskar Thaeter, Christian Grashei, Fabian Gülhan, Reza Nasirigerdeh, Xun Ma, Rui Yan, Hao Chen, S. Kevin Zhou, Nassir Navab, Carolin Mogler, Peter Schüffler
机构
*
Institute of Pathology, Technical University of Munich(慕尼黑工业大学病理研究所)
;
Computer Aided Medical Procedures (CAMP), Technical University of Munich(慕尼黑工业大学计算机辅助医疗程序(CAMP))
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Department of Human Pathology, Juntendo University Graduate School of Medicine(顺天堂大学医学研究生院人体病理学部)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Munich Data Science Institute (MDSI)(慕尼黑数据科学研究所)
TikStance: A Multimodal and Hierarchical Dataset for Multi-target Stance Analysis in TikTok Political Conversations
TikStance:用于TikTok政治对话中多目标立场分析的多模态分层数据集
Yazhi Zhang, Fuqiang Niu, Bowen Zhang
机构
*
School of Artificial Intelligence, Shenzhen Technology University(深圳技术大学人工智能学院)
;
School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学网络空间安全学院)
Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation
通过自我场景增强在多模态大语言模型中强化自我中心空间感知
Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng, Zidong Cao, Lutao Jiang, Zixin Zhang, Huiyu Zhou, Xuming Hu
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Guangxi Zhuang Autonomous Region Information Center(广西壮族自治区信息中心)
;
The Hong Kong University of Science and Technology(香港科技大学)
Triple-Phase Multimodal Knowledge Aggregation Framework for Microbial Keratitis Subtype Diagnosis on Slit-Lamp Photography
基于裂隙灯摄影的微生物性角膜炎亚型诊断三相多模态知识聚合框架
Yiqing Wang, Maria A. Woodward, Ziyun Yang, N. Venkatesh Prajna, Chunming He, Leslie M. Niziol, Mercy Pawar, Ming-Chen Lu, Guillermo Amescua, Rachel Wozniak, Sejal Amin, Abinaya Krishnan, Prabhleen Kochar, Sina Farsiu
机构
*
Department of Biomedical Engineering, Duke University(杜克大学生物医学工程系)
;
Kellogg Eye Center, Department of Ophthalmology and Visual Sciences, University of Michigan(密歇根大学凯洛格眼科中心,眼科学与视觉科学系)
;
Department of Cornea and Refractive Surgery Services, Aravind Eye Care System(阿瓦因眼科医疗系统角膜与屈光手术部)
;
Bascom Palmer Eye Institute, Department of Ophthalmology, University of Miami Miller School of Medicine(迈阿密大学米勒医学院巴斯科姆·帕勒眼科研究所,眼科学系)
;
Flaum Eye Institute, Department of Ophthalmology, University of Rochester Medical Center(罗切斯特大学医学中心弗劳姆眼科研究所,眼科学系)
;
Department of Ophthalmology, Henry Ford Hospital(亨利福特医院眼科部)
;
Duke Eye Center, Duke University School of Medicine(杜克大学医学院杜克眼科中心)
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Pluralis v0.1:迈向用于人工智能风险与可靠性的多元文化、多模态、多语言基准测试
Alicia Parrish, Rajat Shinde, Sanket Badhe, Xinyi Bai, Sree Bhargavi Balija, Hua-Rong Chu, Emilio Ferrara, Armstrong Foundjem, Rajat Ghosh, Aakash Gupta, Xuanli He, Ong Chen Hui, Minji Jung, Madhangi Karimanal, Faiza Khan Khattak, Boryoung Kim, Eugenia Kim, Liliya Lavitas, Seok Min Lim, Victor Lu, Jim Moirangthem, Dhivya Nagasubramanian, Deepak Pandita, Sita Rajagopal, Geetha Raju, Evgeniia Razumovskaia, Aravind Reddy, Federico Ricciuti, Nobin Sarwar, Sungpil Shin, Sunayana Sitaram, Snehal Thorat, Tharindu Cyril Weerasooriya, Jasmijn Bastings, Joachim Baumann, Kongtao Chen, Murali Emani, Mariya Hendriksen, Jiho Jin, Jun Seong Kim, Younghoon Ko, Alicja Kwasniewska, Minjae Lee, Tom Wei-cyuan Lin Kashyap Ramanandula Manjusha, Junho Myung, Junyeong Park, Roma Patel, Shyam Ratan, Sudarsun Santhiappan, Priyanka Suresh, Tuesday, Ksheeraj Sai Vepuri Laura Amortegui-Ordonez, Claire Dennis, Minsuk Kahng, Chris Knotz, Alice Oh, Balaraman Ravindran, Soojung Ryu William Bartholomew, Hiwot Tesfaye, Lora Aroyo
机构
*
Google DeepMind(谷歌DeepMind)
;
University of Alabama in Huntsville(阿拉巴马大学亨茨维尔分校)
;
Google(谷歌)
;
University of Missouri Columbia(密苏里大学哥伦比亚分校)
;
Chunghwa Telecom Laboratories(春木电信实验室)
;
University of Southern California(南加州大学)
;
Polytechnique Montreal(蒙特利尔理工学院)
;
Nutanix
;
ThinkEvolve Labs(ThinkEvolve实验室)
;
UCL(伦敦大学学院)
;
Infocomm Media Development Authority(信息通信媒体发展局)
;
Monark Health(Monark健康)
;
Seoul National University(首尔国立大学)
;
Microsoft(微软)
;
Centre for Responsible AI (CeRAI), Wadhwani School of Data Science and AI (WSAI), Indian Institute of Technology Madras(负责任人工智能中心(CeRAI)、瓦达威人工智能学校(WSAI)、印度理工学院马德拉斯分校)
;
University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
;
Microsoft Research India(微软印度研究院)
;
Stanford University(斯坦福大学)
;
Argonne National Laboratory(阿贡国家实验室)
;
University of Oxford(牛津大学)
;
KAIST(韩国科学技术院)
;
Yonsei University(延世大学)
;
Amazon(亚马逊)
;
UIUC(伊利诺伊大学香槟分校)
;
Rochester Institute of Technology(罗切斯特理工学院)
;
Xenoscube Inc.(Xenoscube公司)
;
Korea AI Safety Institute (K-AISI)(韩国人工智能安全研究所(K-AISI))
;
MLCommons
;
CommonGround
;
Artifex Labs(Artifex实验室)
机构
*
Department of Information Engineering and Computer Science, University of Trento(特伦托大学信息工程与计算机科学系)
;
School of Information and Electronics, Beijing Institute of Technology(北京理工大学信息与电子学院)
;
Beijing Key Laboratory of Fractional Signals and Systems, Beijing Institute of Technology(北京理工大学分数域信号与系统北京市重点实验室)
;
School of Electronic Engineering, Xidian University(西安电子科技大学电子工程学院)
;
Key Laboratory of Collaborative Intelligent Systems of Ministry of Education, Xidian University(西安电子科技大学教育部协同智能系统重点实验室)
;
Academy of Artificial Intelligence, Inner Mongolia Normal University(内蒙古师范大学人工智能研究院)
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
ZeroBench:当代大型多模态模型的一个不可能的视觉基准测试
Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma, Akash Gupta, Samuel Roberts, Ioana Croitoru, Simion-Vlad Bogolin, Jialu Tang, Florian Langer, Vyas Raina, Vatsal Raina, Hanyi Xiong, Vishaal Udandarao, Jingyi Lu, Shiyang Chen, Sam Purkis, Tianshuo Yan, Wenye Lin, Gyungin Shin, Qiaochu Yang, Anh Totti Nguyen, David I. Atkinson, Aaditya Baranwal, Alexandru Coca, Mikah Dang, Sebastian Dziadzio, Jakob D. Kunz, Kaiqu Liang, Alexander Lo, Brian Pulfer, Steven Walton, Charig Yang, Kai Han, Samuel Albanie
机构
*
University of Cambridge(剑桥大学)
;
University of Alberta(阿尔伯塔大学)
;
The University of Hong Kong(香港大学)
;
University of Oxford(牛津大学)
;
Northeastern University(东北大学)
;
Astadeus
;
College of Southern Maryland(马里兰州南部学院)
;
University of Geneva(日内瓦大学)
;
University of Oregon(俄勒冈大学)
Framework and Multi-modal Dataset for Roadwork Zone Detection and Geo-localization
道路施工区域检测与地理定位的框架和多模态数据集
Zhiran Yan, Yutong Xin, S Shyam Shenoi, Rui Song, Gordon Elger
机构
*
Institute of Innovative Mobility (IIMo), Technical University Ingolstadt of Applied Sciences(应用科学英戈尔施塔特技术大学创新移动研究所)
;
Fraunhofer Institute for Transportation and Infrastructure Systems IVI(弗劳恩霍夫交通与基础设施系统研究所IVI)
;
Technical University of Munich(慕尼黑工业大学)
An Automated Multimodal Glaucoma Detection Framework Using ViT and a Stacking-Based Ensemble
使用ViT和基于堆叠的集成方法的自动多模态青光眼检测框架
Ishrat Jahan, Muhammad E. H Chowdhury, Murugappan Murugappan, Kanchon Kanti Podder, Tawsifur Rahman, Shrestha Datta, Md Sakib Abrar Hossain, Md Mosarrof Hossen, Yosra Magdi Salih Mekki, Sanjiban Sekhar Roy
机构
*
Department of Computer Science and Engineering, Shahjalal University of Science and Technology(肖哈尔大学科学与技术学院计算机科学与工程系)
;
Department of Electrical Engineering, Qatar University(卡塔尔大学电气工程系)
;
Department of Electronics and Communication Engineering, Kuwait College of Science and Technology(科威特科学与技术学院电子与通信工程系)
;
Department of Interdisciplinary Engineering, Kennesaw State University(肯尼斯州立大学跨学科工程系)
;
Department of Biomedical Engineering, School of Medicine, Johns Hopkins University(约翰霍普金斯大学医学院生物医学工程系)
;
Department of Biomedical Engineering, University of Oxford(牛津大学生物医学工程系)
;
Department of Computer Science and Engineering, Vellore Institute of Technology(维洛雷理工学院计算机科学与工程系)
TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation
TeachObs:多模态教学观察与模型评估的人工验证基准
Yeil Jeong, Youngjin Yoo, Jiyoung Bae, Seobin Sohn, Hyejin Han, Jinseo Lee, Howard Scott, Unggi Lee
机构
*
Indiana University Bloomington(印第安纳大学布卢明顿分校)
;
Pai Chai University(培才大学)
;
Seoul National University(首尔国立大学)
;
Ewha Womans University(成均馆大学)
;
University of Wolverhampton(沃尔夫汉普顿大学)
;
Korea University Sejong Campus(韩国大学世宗校区)
机构
*
Graduate School of Information Science and Technology, Hokkaido University(北海道大学信息科学研究生院)
;
Graduate School of Computer Science, George Mason University(乔治·马歇尔大学计算机科学研究生院)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
VKnowU:评估多模态语言模型中的视觉知识理解
Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu, Xiangyu Zeng, Limin Wang, Yu Qiao, Yi Wang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Nanjing University(南京大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
City University of Hong Kong(香港城市大学)