ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
ClinFusion:用于整体医学理解的以视觉为中心的多模态大语言模型系统
Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
机构
*
DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)
;
Hupan Laboratory(湖畔实验室)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Department of Radiology, The Affiliated Yangming Hospital of Ningbo University(宁波大学附属阳明医院放射科)
;
Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University(浙江大学伊利诺伊大学厄巴纳香槟校区联合学院,浙江大学)
;
Hepato-Pancreato-Biliary Center, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University(清华长庚医院肝胆胰中心,清华大学临床医学院,清华医学,清华大学)
;
School of Software, Tsinghua University(清华大学软件学院)
;
Beijing National Research Center for Information Science and Technology, Tsinghua University(清华大学北京信息科学与技术国家研究中心)
机构
*
School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学)
;
School of Computer Science and Technology, Northwestern Polytechnical University(计算机科学与技术学院,西北工业大学)
Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation
NLPCC 2026共享任务1概述:难度感知多语言多模态医学教学视频理解评估
Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li
机构
*
School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与工程学院)
;
Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算学系)
;
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
MMR-V:未言明的是什么?视频中多模态深度推理的基准测试
Kejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li, Shangqing Tu, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Tsinghua University(清华大学)
机构
*
Centennial High School, Frisco, Texas, USA(Centennial High School, Texas, USA)
;
Lebanon Trail High School, Frisco, Texas, USA(Lebanon Trail High School, Texas, USA)
;
West Windsor-Plainsboro High School, Princeton Junction, New Jersey, USA(West Windsor-Plainsboro High School, New Jersey, USA)
;
Algoverse AI Research, Palo Alto, California, USA(Algoververse AI Research, California, USA)
Comments14 pages, 4 figures, 8 tables. Presented at the 39th Conference on Neural Information Processing Systems Workshop: VLM4RWD. Presented at the 43th International Conference on Machine Learning Workshops: ICML 2026 CTB, ICML 2026 FAGEN, ICML 2026 EMM-QA. Authors Aahana Basappa and Pranay Goel contributed equally. Code: https://github.com/AahanaB24/AMVICC, Data: https://doi.org/10.5281/zenodo.17646068
An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?
聚焦女性安全分析:多模态数据集能否增强VAD模型?
Sangeeta ., Maddikuntla Sai Prajwal, Debi Prosad Dogra, Kamalakar Vijay Thakare, Hyungjoo Jung, Ig-Jae Kim, Heeseung Choi
机构
*
Indian Institute of Technology Bhubaneswar(印度理工学院巴特那分校)
;
Artificial Intelligence and Robotics Institute, Korea Institute of Science and Technology(人工智能与机器人研究所,韩国科学技术院)
;
Yonsei-KIST Convergence Research Institute, Yonsei University(延世大学KIST融合研究中心)
RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model
RetiBridge:用知识引导的多模态大语言模型连接定量视网膜生物标志物与定性诊断
Zhuangzhi Gao, Hongyi Qin, He Zhao, Qinkai Yu, Feixiang Zhou, Fu Wang, Jinru Ding, Eduard Shantsila, Uazman Alam, Alena Shantsila, Wahbi El-Bouri, Gregory Y. H. Lip, Yalin Zheng
机构
*
University of Liverpool(利物浦大学)
;
Institute of Life Course & Medical Sciences(生命课程与医学科学研究院)
;
Department of Eye and Vision Sciences(眼科与视觉科学系)
;
Computer Science Department(计算机科学系)
;
Cardiovascular & Metabolic Medicine(心血管与代谢医学)
;
Liverpool Centre for Cardiovascular Science(利物浦心血管科学中心)
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
迈向医学数据的视觉-语言基础模型:越南语PET/CT报告生成的多模态数据集和基准
Huu Tien Nguyen, Dac Thai Nguyen, The Minh Duc Nguyen, Trung Thanh Nguyen, Thao Nguyen Truong, Huy Hieu Pham, Johan Barthelemy, Minh Quan Tran, Thanh Tam Nguyen, Quoc Viet Hung Nguyen, Quynh Anh Chau, Hong Son Mai, Thanh Trung Nguyen, Phi Le Nguyen
机构
*
AI4LIFE, Hanoi University of Science and Technology, Vietnam(AI4LIFE,河内科学技术大学,越南)
;
Nagoya University, Japan(名古屋大学,日本)
;
AIST, Japan(日本国家先进工业技术研究院)
;
VinUniversity, Vietnam(文园大学,越南)
;
NVIDIA, USA(NVIDIA,美国)
;
Griffith University, Australia(格里菲斯大学,澳大利亚)
;
Hanoi Medical University, Vietnam(河内医学院,越南)
;
Military Central Hospital, Vietnam(越南108中央军医院)
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(信息多媒体国家重点实验室,计算机学院,北京大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Institute for Brain and Intelligence, Fudan University(脑与智能研究院,复旦大学)
;
University of Science and Technology Beijing(北京科技大学)
;
Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
NeuralBoneReg: An Instance-Specific Label-Free Point Cloud-Based Method for Multi-Modal Bone Surface Registration
NeuralBoneReg:一种用于多模态骨表面注册的实例特定无标签点云方法
Luohong Wu, Matthias Seibold, Nicola A. Cavalcanti, Yunke Ao, Roman Flepp, Aidana Massalimova, Lilian Calvet, Philipp Fürnstahl
机构
*
Research in Orthopedic Computer Science, Balgrist University Hospital, University of Zurich(骨科计算机科学研究所,巴尔格里斯大学医院,苏黎世大学)
;
AI Center, ETH Zurich(人工智能中心,苏黎世联邦理工学院)
机构
*
Department of Automation, Tsinghua University(清华大学自动化系)
;
Institute for Embodied Intelligence and Robotics, Tsinghua University(清华大学具身智能与机器人研究所)
;
TetraBOT Intelligence Co., Ltd.(天博智能科技有限公司)
;
DAMO Academy, Alibaba Group(阿里巴巴达摩院)
;
School of Automation, Southeast University(东南大学自动化学院)
A fine-grained attention and geometric correspondence model for musculoskeletal risk classification in athletes using multimodal visual and skeletal features
机构
*
Department of Computer Science and Engineering, United International University(计算机科学与工程系,国际联合大学)
;
Department of Data Science and Artificial Intelligence, Monash University(数据科学与人工智能系,墨尔本大学)
;
Faculty of Science and Technology, Charles Darwin University(科学与技术学院,查尔斯达尔文大学)
;
Applied Artificial Intelligence and Intelligent Systems (AAIINS) Laboratory, Dhaka(应用人工智能与智能系统实验室,达卡)
Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
用于多模态话语语义发现的无监督多模态聚类
Hanlei Zhang, Hua Xu, Fei Long, Xin Wang, Kai Gao
机构
*
State Key Laboratory of Intelligent Technology and Systems, Department of Computer Science and Technology, Tsinghua University(智能技术与系统国家重点实验室,计算机科学与技术系,清华大学)
;
School of Information Science and Engineering, Hebei University of Science and Technology(信息科学与工程学院,河北科技大学)
;
Samton (Jiangxi) Technology Development Co.,Ltd(江西松通科技发展有限公司)
机构
*
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology (HKUST), Hong Kong SAR, China(香港科技大学计算机科学与工程系)
;
Tencent, Shenzhen, China(腾讯(中国深圳))
;
Shenzhen Institute of Advanced Technology (SIAT), Chinese Academy of Sciences, Shenzhen, China(深圳先进技术研究所(SIAT),中国科学院)