机构
*
Missouri University of Science and Technology(密苏里科技大学)
;
University of South Florida(佛罗里达州立大学)
;
Visa Inc.(Visa公司)
;
George Mason University(乔治·马歇尔大学)
机构
*
Hong Kong Polytechnic University(香港理工大学)
;
Nanyang Technological University(南洋理工大学)
;
Tsinghua University(清华大学)
;
National University of Singapore(新加坡国立大学)
Comments10 pages (36 including references and appendices), 11 figures, accepted at COLM 2026, earlier version accepted at AAAI 2025 Workshop on Document Understanding and Intelligence
机构
*
Beijing Academy of Artificial Intelligence (BAAI), China(北京人工智能研究院)
;
Institute of Information Engineering, Chinese Academy of Sciences, China(信息工程研究所)
;
Beijing University of Technology, China(北京理工大学)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
SenseTime Research(商汤科技研究院)
;
National University of Singapore(新加坡国立大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
ClinFusion:用于整体医学理解的以视觉为中心的多模态大语言模型系统
Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
机构
*
DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)
;
Hupan Laboratory(湖畔实验室)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Department of Radiology, The Affiliated Yangming Hospital of Ningbo University(宁波大学附属阳明医院放射科)
;
Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University(浙江大学伊利诺伊大学厄巴纳香槟校区联合学院,浙江大学)
;
Hepato-Pancreato-Biliary Center, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University(清华长庚医院肝胆胰中心,清华大学临床医学院,清华医学,清华大学)
;
School of Software, Tsinghua University(清华大学软件学院)
;
Beijing National Research Center for Information Science and Technology, Tsinghua University(清华大学北京信息科学与技术国家研究中心)
机构
*
School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学)
;
School of Computer Science and Technology, Northwestern Polytechnical University(计算机科学与技术学院,西北工业大学)
Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation
NLPCC 2026共享任务1概述:难度感知多语言多模态医学教学视频理解评估
Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li
机构
*
School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与工程学院)
;
Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算学系)
;
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
MMR-V:未言明的是什么?视频中多模态深度推理的基准测试
Kejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li, Shangqing Tu, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Tsinghua University(清华大学)
机构
*
Centennial High School, Frisco, Texas, USA(Centennial High School, Texas, USA)
;
Lebanon Trail High School, Frisco, Texas, USA(Lebanon Trail High School, Texas, USA)
;
West Windsor-Plainsboro High School, Princeton Junction, New Jersey, USA(West Windsor-Plainsboro High School, New Jersey, USA)
;
Algoverse AI Research, Palo Alto, California, USA(Algoververse AI Research, California, USA)
Comments14 pages, 4 figures, 8 tables. Presented at the 39th Conference on Neural Information Processing Systems Workshop: VLM4RWD. Presented at the 43th International Conference on Machine Learning Workshops: ICML 2026 CTB, ICML 2026 FAGEN, ICML 2026 EMM-QA. Authors Aahana Basappa and Pranay Goel contributed equally. Code: https://github.com/AahanaB24/AMVICC, Data: https://doi.org/10.5281/zenodo.17646068
An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?
聚焦女性安全分析:多模态数据集能否增强VAD模型?
Sangeeta ., Maddikuntla Sai Prajwal, Debi Prosad Dogra, Kamalakar Vijay Thakare, Hyungjoo Jung, Ig-Jae Kim, Heeseung Choi
机构
*
Indian Institute of Technology Bhubaneswar(印度理工学院巴特那分校)
;
Artificial Intelligence and Robotics Institute, Korea Institute of Science and Technology(人工智能与机器人研究所,韩国科学技术院)
;
Yonsei-KIST Convergence Research Institute, Yonsei University(延世大学KIST融合研究中心)
RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model
RetiBridge:用知识引导的多模态大语言模型连接定量视网膜生物标志物与定性诊断
Zhuangzhi Gao, Hongyi Qin, He Zhao, Qinkai Yu, Feixiang Zhou, Fu Wang, Jinru Ding, Eduard Shantsila, Uazman Alam, Alena Shantsila, Wahbi El-Bouri, Gregory Y. H. Lip, Yalin Zheng
机构
*
University of Liverpool(利物浦大学)
;
Institute of Life Course & Medical Sciences(生命课程与医学科学研究院)
;
Department of Eye and Vision Sciences(眼科与视觉科学系)
;
Computer Science Department(计算机科学系)
;
Cardiovascular & Metabolic Medicine(心血管与代谢医学)
;
Liverpool Centre for Cardiovascular Science(利物浦心血管科学中心)
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
迈向医学数据的视觉-语言基础模型:越南语PET/CT报告生成的多模态数据集和基准
Huu Tien Nguyen, Dac Thai Nguyen, The Minh Duc Nguyen, Trung Thanh Nguyen, Thao Nguyen Truong, Huy Hieu Pham, Johan Barthelemy, Minh Quan Tran, Thanh Tam Nguyen, Quoc Viet Hung Nguyen, Quynh Anh Chau, Hong Son Mai, Thanh Trung Nguyen, Phi Le Nguyen
机构
*
AI4LIFE, Hanoi University of Science and Technology, Vietnam(AI4LIFE,河内科学技术大学,越南)
;
Nagoya University, Japan(名古屋大学,日本)
;
AIST, Japan(日本国家先进工业技术研究院)
;
VinUniversity, Vietnam(文园大学,越南)
;
NVIDIA, USA(NVIDIA,美国)
;
Griffith University, Australia(格里菲斯大学,澳大利亚)
;
Hanoi Medical University, Vietnam(河内医学院,越南)
;
Military Central Hospital, Vietnam(越南108中央军医院)
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(信息多媒体国家重点实验室,计算机学院,北京大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Institute for Brain and Intelligence, Fudan University(脑与智能研究院,复旦大学)
;
University of Science and Technology Beijing(北京科技大学)
;
Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
NeuralBoneReg: An Instance-Specific Label-Free Point Cloud-Based Method for Multi-Modal Bone Surface Registration
NeuralBoneReg:一种用于多模态骨表面注册的实例特定无标签点云方法
Luohong Wu, Matthias Seibold, Nicola A. Cavalcanti, Yunke Ao, Roman Flepp, Aidana Massalimova, Lilian Calvet, Philipp Fürnstahl
机构
*
Research in Orthopedic Computer Science, Balgrist University Hospital, University of Zurich(骨科计算机科学研究所,巴尔格里斯大学医院,苏黎世大学)
;
AI Center, ETH Zurich(人工智能中心,苏黎世联邦理工学院)