DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
Zijie Meng, Jin Hao, Xiwei Dai, Yang Feng, Jiaxiang Liu, Bin Feng, Huikai Wu, Xiaotang Gai, Hengchuan Zhu, Tianxiang Hu, Yangyang Wu, Hongxia Xu, Jin Li, Jun Xiao, Xiaoqiang Liu, Joey Tianyi Zhou, Fudong Zhu, Zhihe Zhao, Lunguo Xia, Bing Fang, Jimeng Sun, Jian Wu, Zuozhu Liu
机构
*
Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang University, Hangzhou(牙科医院,口腔医学院,浙江大学医学院,浙江大学,杭州)
;
College of Computer Science and Technology, Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University, Hangzhou(计算机科学与技术学院,浙江大学-伊利诺伊大学 Urbana-Champaign 院,浙江大学,杭州)
;
Department of Orthodontics, Shanghai Ninth People’s Hospital, College of Stomatology, Shanghai Jiao Tong University School of Medicine(正畸科,上海第九人民医院,口腔医学院,上海交通大学医学院)
;
Angelalign Technology Inc.(Angelalign 技术公司)
机构
*
State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)
;
Chongqing University of Posts and Telecommunications(重庆邮电大学)
True Multimodal In-Context Learning Needs Attention to the Visual Context
Shuo Chen, Jianzhe Liu, Zhen Han, Yan Xia, Daniel Cremers, Philip Torr, Volker Tresp, Jindong Gu
机构
*
LMU Munich(慕尼黑大学)
;
Technical University of Munich(慕尼黑技术大学)
;
Siemens AG(西门子股份公司)
;
University of Science and Technology of China(中国科学技术大学)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Konrad Zuse School of Excellence in Reliable AI (relAI)(Konrad Zuse可靠性人工智能卓越学院)
;
University of Oxford(牛津大学)
Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
Junyu Chen, Yihua Gao, Mingyuan Ge, Mingyong Li
机构
*
College of Computer and Information Science, Chongqing Normal University(重庆师范大学计算机与信息科学学院)
;
School of Big Data and Software Engineering, Chongqing University(重庆大学大数据与软件工程学院)
机构
*
Hong Kong Baptist University(香港 Baptist 大学)
;
Shanxi University(山西大学)
;
Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海研究院)
;
Tencent AI Lab(腾讯人工智能实验室)
;
Manchester Metropolitan University(曼彻斯特 Metropolitan 大学)
;
The University of Manchester(曼彻斯特大学)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
Yufei Zhan, Hongyin Zhao, Yousong Zhu, Shurong Zheng, Fan Yang, Ming Tang, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Peng Cheng Laboratory, Shenzhen, China(鹏城实验室)
;
Wuhan AI Research, Wuhan, China(武汉人工智能研究所)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
Yue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta, Luca Weihs, Andrew Head, Mark Yatskar, Chris Callison-Burch, Ranjay Krishna, Aniruddha Kembhavi, Christopher Clark
机构
*
University of Pennsylvania(宾夕法尼亚大学)
;
Allen Institute for Artificial Intelligence(人工智能研究所)