Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation
NLPCC 2026共享任务1概述:难度感知多语言多模态医学教学视频理解评估
Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li
机构
*
School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与工程学院)
;
Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算学系)
;
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
MMR-V:未言明的是什么?视频中多模态深度推理的基准测试
Kejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li, Shangqing Tu, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Tsinghua University(清华大学)
The In-Car Sign Language Corpus (ICSL): A Multi-Modal Resource for Constrained-Space Sign Language Recognition
车内手语语料库(ICSL):用于受限空间手语识别的多模态资源
Raviteja Boddu, Guilherme Vieira Leite, Joed Lopes da Silva, Ângelo Benetti, Isabela Barbieri, Natália de Melo Afonso, Thyago Santos, Helio Pedrini, Felipe Venâncio Barbosa, José Mario De Martino, Munir Georges, Alessandro Zimmer
机构
*
Technische Hochschule Ingolstadt (THI)(因戈尔施塔特技术大学)
;
Universidade Estadual de Campinas (UNICAMP)(坎皮纳斯州立大学)
;
Universidade de São Paulo (USP)(圣保罗大学)
机构
*
Centennial High School, Frisco, Texas, USA(Centennial High School, Texas, USA)
;
Lebanon Trail High School, Frisco, Texas, USA(Lebanon Trail High School, Texas, USA)
;
West Windsor-Plainsboro High School, Princeton Junction, New Jersey, USA(West Windsor-Plainsboro High School, New Jersey, USA)
;
Algoverse AI Research, Palo Alto, California, USA(Algoververse AI Research, California, USA)
Comments14 pages, 4 figures, 8 tables. Presented at the 39th Conference on Neural Information Processing Systems Workshop: VLM4RWD. Presented at the 43th International Conference on Machine Learning Workshops: ICML 2026 CTB, ICML 2026 FAGEN, ICML 2026 EMM-QA. Authors Aahana Basappa and Pranay Goel contributed equally. Code: https://github.com/AahanaB24/AMVICC, Data: https://doi.org/10.5281/zenodo.17646068
A Quantitative Analysis of Multimodal Biomarkers in Alzheimer's Disease
阿尔茨海默病多模态生物标志物的定量分析
Antonio Scardace, Daniele Ravì
机构
*
Department of Mathematics and Computer Science(数学与计算机科学系)
;
University of Catania(卡塔尼亚大学)
;
Department MIFT(MIFT部门)
;
University of Messina(梅西纳大学)
机构
*
School of Electronic and Electrical Engineering, Shanghai University of Engineering Science(上海工程技术大学电子与电气工程学院)
;
Tencent Youtu Lab(腾讯云视频实验室)
;
ENT Institute and Department of Otorhinolaryngology, Eye & ENT Hospital of Fudan University(复旦大学耳鼻喉科医院耳鼻喉科研究所)
;
National University of Singapore(新加坡国立大学)
专题命中
多模态评测
:multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI
EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models
EmoBench-M:多模态大语言模型情感智能评估基准
He Hu, Lianzhong You, Hongbo Xu, Qianning Wang, Fei Richard Yu, Fei Ma, Zebang Cheng, Zheng Lian, Yucheng Zhou, Laizhong Cui
机构
*
Shenzhen University(深圳大学)
;
Guangdong Laboratory of Artificial Intelligence(广东人工智能与数字经济实验室)
;
Auckland University of Technology(奥克兰理工大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
SKL-IOTSC, CIS, University of Macau(澳门科学技术大学SKL-IOTSC、CIS)
PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models
PRISM-XR: 通过多模态大语言模型赋能隐私感知的扩展现实协作
Jiangong Chen, Mingyu Zhu, Bin Li
机构
*
Department of Electrical Engineering, The Pennsylvania State University, University Park, PA 16802, USA(电气工程系,宾夕法尼亚州立大学,University Park,PA 16802,USA)
机构
*
National Key Laboratory for Multimedia Information Processing(多媒体信息处理国家重点实验室)
;
Peking University School of Computer Science(北京大学计算机科学学院)
;
Peking University School of Software and Microelectronics(北京大学软件与微电子学院)
MuDD: A Multimodal Deception Detection Dataset and GSR-Guided Progressive Distillation for Non-Contact Deception Detection
MuDD:一种多模态欺骗检测数据集和GSR引导的渐进性知识蒸馏用于非接触欺骗检测
Peiyuan Jiang, Yao Liu, Yanglei Gan, Jiaye Yang, Lu Liu, Daibing Yao, Qiao Liu
机构
*
School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
;
School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
机构
*
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
University of Science and Technology of China(中国科学技术大学)
;
Arizona State University(亚利桑那州立大学)
;
Honda Research Institute, USA(本田美国研究所)
FLEX: A Largescale Multimodal, Multiview Dataset for Learning Structured Representations for Fitness Action Quality Assessment
FLEX:一个大规模多模态、多视角数据集,用于学习结构化表示以评估健身动作质量
Hao Yin, Lijun Gu, Paritosh Parmar, Lin Xu, Tianxiao Guo, Xiujin Liu, Weiwei Fu, Yang Zhang, Tianyou Zheng
机构
*
School of Biomedical Engineering (Suzhou), USTC(中国科学技术大学苏州生物医学工程学院)
;
Suzhou Institute of Biomedical Engineering and Technology, CAS(中国科学院苏州生物医学工程技术研究所)
;
Institute of High-Performance Computing, A*STAR, Singapore(新加坡科技研究局高性能计算研究所)
;
School of Psychology, Beijing Sports University(北京体育大学心理学院)
;
School of Competitive Sports, Beijing Sports University(北京体育大学竞技体育学院)
;
Department of Robotics, University of Michigan(密歇根大学机器人系)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Casivision(中科视语)
;
Weiqiao-UCAS Science and Technology Park(魏桥国科科技园)