机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(信息多媒体国家重点实验室,计算机学院,北京大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Institute for Brain and Intelligence, Fudan University(脑与智能研究院,复旦大学)
;
University of Science and Technology Beijing(北京科技大学)
;
Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
Explainable AI for Screening Abuse-Related Trauma in Bangladeshi Children: A Training-Free Multimodal Framework Evaluated on Noise-Aware Synthetic Data
机构
*
Southwestern University of Finance and Economics(西南财经大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Central South University(中南大学)
;
Hithink Research(Hithink研究)
;
Westlake University(西湖大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
University of Manchester(曼彻斯特大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
University of Adelaide(阿德莱德大学)
;
Fudan University(复旦大学)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
Chengdu Everimaging Science and Technology Co., Ltd(成都亿联科技有限公司)
GestaltMML: Enhancing Rare Genetic Disease Diagnosis through Multimodal Machine Learning Combining Facial Images and Clinical Text
GestaltMML:通过结合面部图像和临床文本的多模态机器学习增强罕见遗传病诊断
Da Wu, Zhanliang Wang, Hongzhuo Chen, Jingye Yang, Cong Liu, Tzung-Chien Hsieh, Elaine Marchi, Justin Blair, Peter Krawitz, Chunhua Weng, Wendy Chung, Gholson J. Lyon, Ian D. Krantz, Jennifer M. Kalish, Kai Wang
机构
*
Raymond G. Perelman Center for Cellular and Molecular Therapeutics, Children’s Hospital of Philadelphia(雷蒙德·G·佩尔曼细胞与分子治疗中心,费城儿童医院)
;
Department of Mathematics, University of Pennsylvania(数学系,宾夕法尼亚大学)
;
Department of Biomedical Informatics, Columbia University Irving Medical Center(生物医学信息学系,哥伦比亚大学伊万斯医疗中心)
;
Department of Human Genetics, New York State Institute for Basic Research in Developmental Disabilities, Staten Island, NY, USA(人类遗传学系,纽约州发育障碍基础研究机构,纽约州史泰登岛)
;
Division of Human Genetics, Children’s Hospital of Philadelphia(人类遗传学部,费城儿童医院)
;
Department of Pediatrics, Boston Children’s Hospital, Harvard Medical School(儿科系,波士顿儿童医院,哈佛医学院)
;
Biology PhD Program, The Graduate Center, The City University of New York(生物学博士项目,纽约市立大学研究生中心)
;
Department of Genetics, Perelman School of Medicine, University of Pennsylvania(遗传学系,宾夕法尼亚大学佩尔曼医学学院)
;
Department of Pediatrics, Perelman School of Medicine, University of Pennsylvania(儿科系,宾夕法尼亚大学佩尔曼医学学院)
;
Department of Pathology and Laboratory Medicine, Perelman School of Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学佩尔曼医学学院)
Framework and Multi-modal Dataset for Roadwork Zone Detection and Geo-localization
道路施工区域检测与地理定位的框架和多模态数据集
Zhiran Yan, Yutong Xin, S Shyam Shenoi, Rui Song, Gordon Elger
机构
*
Institute of Innovative Mobility (IIMo), Technical University Ingolstadt of Applied Sciences(应用科学英戈尔施塔特技术大学创新移动研究所)
;
Fraunhofer Institute for Transportation and Infrastructure Systems IVI(弗劳恩霍夫交通与基础设施系统研究所IVI)
;
Technical University of Munich(慕尼黑工业大学)
An Automated Multimodal Glaucoma Detection Framework Using ViT and a Stacking-Based Ensemble
使用ViT和基于堆叠的集成方法的自动多模态青光眼检测框架
Ishrat Jahan, Muhammad E. H Chowdhury, Murugappan Murugappan, Kanchon Kanti Podder, Tawsifur Rahman, Shrestha Datta, Md Sakib Abrar Hossain, Md Mosarrof Hossen, Yosra Magdi Salih Mekki, Sanjiban Sekhar Roy
机构
*
Department of Computer Science and Engineering, Shahjalal University of Science and Technology(肖哈尔大学科学与技术学院计算机科学与工程系)
;
Department of Electrical Engineering, Qatar University(卡塔尔大学电气工程系)
;
Department of Electronics and Communication Engineering, Kuwait College of Science and Technology(科威特科学与技术学院电子与通信工程系)
;
Department of Interdisciplinary Engineering, Kennesaw State University(肯尼斯州立大学跨学科工程系)
;
Department of Biomedical Engineering, School of Medicine, Johns Hopkins University(约翰霍普金斯大学医学院生物医学工程系)
;
Department of Biomedical Engineering, University of Oxford(牛津大学生物医学工程系)
;
Department of Computer Science and Engineering, Vellore Institute of Technology(维洛雷理工学院计算机科学与工程系)
TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation
TeachObs:多模态教学观察与模型评估的人工验证基准
Yeil Jeong, Youngjin Yoo, Jiyoung Bae, Seobin Sohn, Hyejin Han, Jinseo Lee, Howard Scott, Unggi Lee
机构
*
Indiana University Bloomington(印第安纳大学布卢明顿分校)
;
Pai Chai University(培才大学)
;
Seoul National University(首尔国立大学)
;
Ewha Womans University(成均馆大学)
;
University of Wolverhampton(沃尔夫汉普顿大学)
;
Korea University Sejong Campus(韩国大学世宗校区)
机构
*
Graduate School of Information Science and Technology, Hokkaido University(北海道大学信息科学研究生院)
;
Graduate School of Computer Science, George Mason University(乔治·马歇尔大学计算机科学研究生院)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
VKnowU:评估多模态语言模型中的视觉知识理解
Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu, Xiangyu Zeng, Limin Wang, Yu Qiao, Yi Wang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Nanjing University(南京大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
City University of Hong Kong(香港城市大学)
EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions
EmoteGPT:从自然语言描述生成3D人类面部表情
Haoran Wang, Mohit Mendiratta, Christian Theobalt, Adam Kortylewski
机构
*
Max Planck Institute for Informatics, Saarland Informatics Campus(马克斯·普朗克信息研究所,萨尔兰信息学园区)
;
CISPA Helmholtz Center for Information Security(亥姆霍兹信息安全中心)
机构
*
Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
University of Manchester(曼彻斯特大学)
;
Institute of Metal Research, Chinese Academy of Sciences(中国科学院金属研究所)
;
China University of Geosciences(中国地质大学)
机构
*
Beihang University(北京航空航天大学)
;
Zhongguancun Academy(中关村科学城)
;
Communication University of China(中国传媒大学)
;
King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
机构
*
School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)
;
Institute of Plasma Physics, Chinese Academy of Sciences(中国科学院等离子体物理研究所)
;
State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, Anhui University(安徽大学光电子信息获取与控制技术国家重点实验室)
HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding
HCSU:用于细粒度历史书法风格理解的数据集和基准测试
Yinsheng Yao, Yan Liu, Chen Ye
机构
*
School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
;
The Key Laboratory of Embedded System and Service Computing, Ministry of Education(教育部嵌入式系统与服务计算重点实验室)
Transition Information Density: Morphological Trajectories, Synesthetic Perception, and Structured Interpolation in Neural Training (or: The Synesthetic AI)
过渡信息密度:神经训练中的形态轨迹、联觉感知和结构化插值(或:联觉人工智能)
Sam Mao
机构
*
New York University(纽约大学)
;
Interactive Media Arts(互动媒体艺术)
Comments38 pages, 9 figures, 4 tables. Empirical results from structured interpolation training across four representational mediums. Pipeline scripts, experimental data, and the Synesthesia Grid algorithm available upon reasonable request