SurgFusion-Net: Diversified Adaptive Multimodal Fusion Network for Surgical Skill Assessment
SurgFusion-Net:用于外科技能评估的多样化自适应多模态融合网络
Runlong He, Freweini M. Tesfai, Matthew W. E. Boal, Nazir Sirajudeen, Dimitrios Anastasiou, Jialang Xu, Mobarak I. Hoque, Philip J. Edwards, John D. Kelly, Ashwin Sridhar, Abdolrahim Kadkhodamohammadi, Dhivya Chandrasekaran, Matthew J. Clarkson, Danail Stoyanov, Nader Francis, Evangelos B. Mazomenos
机构
*
UCL Hawkes Institute and the Department of Medical Physics & Biomedical Engineering, UCL(UCL哈维斯研究所及UCL医学物理与生物医学工程系)
;
UCL Hawkes Institute and the Department of Computer Science, UCL(UCL哈维斯研究所及UCL计算机科学系)
;
UCL Hawkes Institute and the Division of Informatics, Imaging & Data Sciences, The University of Manchester(UCL哈维斯研究所及信息学、成像与数据科学系,曼彻斯特大学)
;
UCL Hospitals NHS Foundation Trust(UCL医院 NHS基金会信托)
;
Griffin Institute, Northwick Park and St Mark’s Hospital(格里芬研究所,北wick公园及圣马可医院)
MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning
MMR-Life: 组合真实场景以进行多模态多图像推理
Jiachun Li, Shaoping Huang, Zhuoran Jin, Chenlong Zhang, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所认知与决策智能复杂系统重点实验室)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学交叉学科学院)
SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
SportR:多模态大语言模型在体育中的推理基准
Haotian Xia, Haonan Ge, Junbo Zou, Hyun Woo Choi, Xuebin Zhang, Danny Suradja, Botao Rui, Ethan Tran, Wendy Jin, Zhen Ye, Xiyang Lin, Christopher Lai, Shengjie Zhang, Junwen Miao, Shichao Chen, Rhys Tracy, Vicente Ordonez, Weining Shen, Hanjie Chen
机构
*
Department of Computer Science, Rice University(Rice大学计算机科学系)
;
Ken Kennedy Institute, Rice University(Rice大学肯尼迪研究所)
;
Department of Statistics, University of California, Irvine(伊利诺伊大学欧文分校统计系)
;
College of Sciences, Georgia Institute of Technology(佐治亚理工学院科学学院)
;
Department of Applied Mathematics and Statistics, Johns Hopkins University(约翰霍普金斯大学应用数学与统计学系)
;
Department of Computer Science, University of California, Santa Barbara(加州大学圣芭芭拉分校计算机科学系)
机构
*
University of Texas(德克萨斯大学)
;
Dell Children’s Medical Center(德尔儿童医学中心)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
University of Nevada, Reno(内华达大学里诺分校)
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
AgentVista: 评估在超挑战性现实视觉场景中的多模态代理
Zhaochen Su, Jincheng Gao, Hangyu Guo, Zhenhua Liu, Lueyang Zhang, Xinyu Geng, Shijue Huang, Peng Xia, Guanyu Jiang, Cheng Wang, Yue Zhang, Yi R. Fung, Junxian He
机构
*
Hong Kong University of Science and Technology(香港理工大学)
;
Zhejiang University(浙江大学)
;
National University of Singapore(新加坡国立大学)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
Yuxuan Yang, Zhonghao Yan, Yi Zhang, Bo Yun, Muxi Diao, Guowei Zhao, Kongming Liang, Wenbin Li, Zhanyu Ma
机构
*
School of Artificial Intelligence, Beijing University of Posts and Telecommunications(人工智能学院,北京邮电大学)
;
Department of Pathology, National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(pathology department, 国家癌症中心/国家癌症临床研究中心/癌症医院, 中国医学科学院和北京协和医学院)
Chenggang Rong, Tao Han, Zhiyuan Zhao, Yaowu Fan, Jia Wan, Song Guo, Yuan Yuan, Junyu Gao
机构
*
Northwestern Polytechnical University(北华大学)
;
Hong Kong University of Science and Technology(香港科技大学)
;
Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究所)
;
Sun Yat-sen University(中山大学)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Shanghai Jiao Tong University(上海交通大学)
;
East China Normal University(华东师范大学)
;
Shanghai University of Electric Power(上海电力大学)
ICYM2I: The illusion of multimodal informativeness under missingness
ICYM2I: 多模态信息性在缺失情况下的错觉
Young Sang Choi, Vincent Jeanselme, Pierre Elias, Shalmali Joshi
机构
*
Department of Biomedical Informatics, Columbia University(生物医学信息学系,哥伦比亚大学)
;
Seymour, Paul, and Gloria Milstein Division of Cardiology, Department of Medicine, Columbia University Irving Medical Center(塞缪尔、保罗和格洛丽亚米尔斯坦心脏病科,哥伦比亚大学伊万格琳医学中心)
Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
为MLLMs定制视觉情绪评估:一种开放词汇、多维且可扩展的方法
Daiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma, Yu Zhou
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
VCIP & TMCC & DISSec, College of Computer Science, Nankai University(南开大学计算机学院)
;
Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理学与认知科学系)
;
University of Chinese Academy of Sciences(中国科学院大学)
FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality Localization
FiLo:通过细粒度描述和高质量定位实现零样本异常检测
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen, Hao Li, Ming Tang, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所基础模型研究中心)
;
University of Chinese Academy of Sciences, Beijing, China(中国科学院大学)
;
Objecteye Inc., Beijing, China(Objecteye公司)
;
Central South University, Hunan, China(中南大学)