CommentsAccepted at the 43rd International Conference on Machine Learning (ICML 2026) Workshop on Efficient Multimodal Question Answering (EMM-QA), Seoul, South Korea. Copyright 2026 by the author(s). (Archival)
Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI
多模态3D MRI中的局部、全局和跨模态上下文建模
Minh Duc Do, Tillmann Rheude, Noel Kronenberg, Roland Eils, Benjamin Wild
机构
*
Berlin Institute of Health at Charité - Universitätsmedizin Berlin(柏林健康研究所,柏林夏里特医学院)
;
Health Data Science Unit, Heidelberg University Hospital and BioQuant(海德堡大学医院与BioQuant健康数据科学部)
;
Intelligent Medicine Institute, Fudan University(复旦大学智能医学研究所)
;
Department of Mathematics and Computer Science, Freie Universität Berlin(柏林自由大学数学与计算机科学系)
机构
*
Research Institute of Trustworthy Autonomous Systems and Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学可信自主系统研究院与计算机科学与工程系)
;
Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)
A Unified Multi-Modal Framework for Intelligent Financial Systems: Integrating Reinforcement Learning, High-Frequency Trading, and Game-Theoretic Approaches with Cross-Modal Sentiment Analysis
面向智能金融系统的统一多模态框架:整合强化学习、高频交易和博弈论方法与跨模态情感分析
Fanrong Liu, Zhang Yuwei, Mingni Luo
机构
*
Henan University, International Eurasia College(河南大学,国际欧亚学院)
;
City University of Hong Kong, College of Business(香港城市大学,商学院)
;
Northeastern University, School of Electronic and Information Engineering(东北大学,电子与信息工程学院)
Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning
用于联邦多模态大语言模型微调的弹性正则化和合成重放持续学习
Jing Liu, Chenxuanyin Zou, Jiayang Ren, Gaoyun Fang, Chengfang Li, Yan Wang, Zhenchao Ma, Bo Hu
机构
*
The University of British Columbia(英属哥伦比亚大学)
;
Fudan University(复旦大学)
;
Royal College of Science, Imperial College London(伦敦帝国理工学院皇家科学学院)
;
Dyson School of Design Engineering(戴森设计工程学院)
;
Suzhou Institute of Biomedical Engineering and Technology (SIBET), Chinese Academy of Sciences(中国科学院苏州生物医学工程技术研究所)
;
East China Normal University(华东师范大学)
Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
关注、变换或静默:面向高效多模态大语言模型推理的算子级视觉跳跃
Zhaoyang Luo, Runmin Dong, Miao Yang, Fan Wei, Yushan Lai, Bin Luo, Haohuan Fu
机构
*
Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
;
Sun Yat-sen University(中山大学)
;
National Supercomputing Center in Shenzhen(国家超级计算深圳中心)
;
Tsinghua University(清华大学)
机构
*
National University of Singapore(新加坡国立大学)
;
Tsinghua University(清华大学)
;
University of Science and Technology of China(中国科学技术大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
University of California, Berkeley(加州大学伯克利分校)
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
TOPS:通过构建令牌最优保留集实现高效多模态大语言模型推理的第一性原理视觉令牌剪枝
Tinghao Wang, Yichen Guo, Rui Huang, Zheng Lu, Qizhe Zhang, Chenxi Li, Yuan Zhang, Jiajun Cao, Zhirong Shen, Yaosong Du, Guangyan Gan, Wenya Wang, Lin William Cong, Shanghang Zhang
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Nanyang Technological University(南洋理工大学)
;
Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)