From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
从结构到协同:多模态大语言模型中视觉-语言感知范式演进综述
Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao, Jiancheng Lv
机构
*
School of Computer Science, Sichuan University(四川大学计算机学院)
;
School of Electronic and Computer Engineering, Peking University Shenzhen Graduate School(北京大学深圳研究生院电子与计算机工程学院)
;
Institute of Artificial Intelligence (TeleAI), China Telecom and Northwestern Polytechnical University(中国电信与西北工业大学人工智能研究院(TeleAI))
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
在自进化大型多模态模型中更加关注视觉标记
Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Salman Khan, Fahad Khan
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Aalto University(阿尔托大学)
;
Australian National University(澳大利亚国立大学)
;
Linköping University(林雪平大学)
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
ReasonCLIP-58M: CLIP的视觉基础常识推理监督
Sicheng Zhang, Muzammal Naseer, Binzhu Xie, Naufal Suryanto, Shi Qiu, Jamal Bentahar, Naveed Akhtar, Mubarak Shah
机构
*
Khalifa University(卡利法大学)
;
University of Western Australia(西澳大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Melbourne(墨尔本大学)
;
University of Central Florida(佛罗里达中央大学)
Through the Looking Glass: A Dual Perspective on Weakly-Supervised Few-Shot Segmentation
透过镜中奇境:弱监督小样本分割的双重视角
Jiaqi Ma, Guo-Sen Xie, Fang Zhao, Zechao Li
机构
*
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
;
School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs
CheXanatomy: 面向胸部X光片的解剖感知视觉-语言建模
Sergios Gatidis, Curtis Langlotz, Christian Bluethgen
机构
*
Stanford Center for Artificial Intelligence in Medicine and Imaging, Stanford University(斯坦福大学医学与影像人工智能中心)
;
Department of Radiology, Stanford University(斯坦福大学放射学系)
机构
*
School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China (USTC)(生物医学工程学院,生命科学与医学系,中国科学技术大学)
;
Center for Medical Imaging, Robotics, Analytic Computing & Learning (MIRACLE)(医学影像、机器人、分析计算与学习中心)
;
Suzhou Institute for Advanced Research, USTC(苏州先进研究院,中国科学技术大学)
;
Department of Radiology, The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, USTC(放射科,中国科学技术大学第一附属医院,生命科学与医学系,中国科学技术大学)
;
T Magnetic Resonance Translational Medicine Research Center, Department of Radiology, The First Affiliated Hospital (Southwest Hospital) of Army Medical University(7T磁共振转化医学研究中心,放射科,中国医学大学第一附属医院(西南医院))
6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
6根手指,1个肾脏:自然对抗性医学图像揭示视觉语言模型的关键弱点
Leon Mayer, Piotr Kalinowski, Caroline Ebersbach, Marcel Knopp, Tim Rädsch, Evangelia Christodoulou, Annika Reinke, Fiona R. Kolbinger, Lena Maier-Hein
机构
*
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems(德国癌症研究中心(DKFZ)海德堡,智能医学系统部门)
;
Medical Faculty, Heidelberg University(海德堡大学医学院)
;
Faculty of Mathematics and Computer Science, Heidelberg University(海德堡大学数学与计算机科学学院)
;
HIDSS4Health - Helmholtz Information and Data Science School for Health, Karlsruhe/Heidelberg(HIDSS4Health - 哈勃-马克斯信息与数据科学健康学院,卡尔斯鲁厄/海德堡)
;
Helmholtz Imaging, German Cancer Research Center (DKFZ)(哈勃-马克斯成像,德国癌症研究中心(DKFZ))
;
Engineering Faculty, Heidelberg University(海德堡大学工程学院)
;
School of Computation, Information and Technology, TUM(技术大学(TUM)计算、信息与技术学院)
;
Weldon School of Biomedical Engineering, Purdue University(普渡大学韦尔登生物医学工程学院)
;
Department of Visceral, Thoracic and Vascular Surgery, University Hospital and Faculty of Medicine Carl Gustav Carus, TUD Dresden University of Technology(visceral、胸腔和血管外科部门,技术大学(TUD)德累斯顿大学医院和医学院)
;
National Center for Tumor Diseases (NCT), NCT Heidelberg, a partnership between DKFZ and University Hospital Heidelberg(肿瘤疾病国家中心(NCT),海德堡NCT,DKFZ与海德堡大学医院之间的合作)
;
Heidelberg University Hospital, Surgical Clinic, Surgical AI Research Group(海德堡大学医院,外科诊所,外科人工智能研究组)
;
Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE(Mohamed Bin Zayed人工智能大学(MBZUAI),阿布扎赫,阿拉伯联合酋长国)
机构
*
Department of Artificial Intelligence, School of Informatics, Xiamen University(厦门大学信息学院人工智能系)
;
Department of Computer Science, Aberystwyth University(阿伯里斯特威斯大学计算机科学系)
Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars
低资源场景下尼泊尔口语词汇到情感条件手语虚拟人物的多模态翻译
Jatin Bhusal, Salma Tamang
机构
*
Center for Human Mobility and Communications, Prateek Innovations(普拉蒂克创新公司人类移动与通信中心)
;
Sunway International Business School, Birmingham City University(双威国际商学院,伯明翰城市大学)
Comments6 pages, 2 figures. Accepted at the 2026 International Joint Conference on Neural Networks (IJCNN 2026), IEEE WCCI 2026; presented as an oral talk. Code and ART-SafeBench benchmark: https://github.com/FujitsuResearch/mirror
A Systematic Survey of Semantic Role Labeling in the Era of Pretrained Language Models
预训练语言模型时代语义角色标注的系统综述
Huiyao Chen, Meishan Zhang, Jing Li, Lilja Øvrelid, Jan Hajič, Hao Fei, Min Zhang
机构
*
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Shenzhen Loop Area Institute (SLAI)(深圳南山区研究院)
;
University of Oslo(奥斯陆大学)
;
Charles University(查尔斯大学)
AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired
AIDEN:面向视障人士的AI助手设计与初步研究
Luis Marquez-Carpintero, Francisco Gomez-Donoso, Zuria Bauer, Bessie Dominguez-Dager, Alvaro Belmonte-Baeza, Mónica Pina-Navarro, Francisco Morillas-Espejo, Felix Escalona, Miguel Cazorla
机构
*
Institute for Computer Research, University of Alicante(计算机研究所,阿利坎特大学)
;
ETH Zurich(苏黎世联邦理工学院)
机构
*
Central South University(中南大学)
;
Tsinghua University(清华大学)
;
South China Normal University(华南师范大学)
;
ByteDance Inc(字节跳动公司)
;
University of Zaragoza(阿拉维达大学)
;
CosmosMind
;
Wuhan University(武汉大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
Southeast University(东南大学)
;
Tencent(腾讯公司)
;
Nankai University(南开大学)
;
Supermicro Computer Inc(Supermicro计算机公司)
;
Huazhong University of Science and Technology(华中科技大学)
Active Adversarial Perturbation-driven Associative Memory Retrieval for RGB-Event Visual Object Tracking
主动对抗扰动驱动的关联记忆检索用于RGB-事件视觉目标跟踪
Xiao Wang, Xufeng Lou, Zikang Yan, Lan Chen, Sibao Chen, Yaowei Wang, Yonghong Tian, Jin Tang
机构
*
School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)
;
School of Electronic and Information Engineering, Anhui University(安徽大学电子信息工程学院)
;
Peng Cheng Laboratory(鹏城实验室)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
School of Computer Science, Peking University(北京大学计算机学院)
;
School of Electronic and Computer Engineering, Shenzhen Graduate School, Peking University(北京大学深圳研究生院电子与计算机工程学院)
机构
*
Shandong Key Laboratory of Ubiquitous Intelligent Computing, School of Information Science and Engineering, University of Jinan(山东省 Ubiquitous Intelligent Computing 重点实验室,济南大学信息科学与工程学院)
;
College of Information Science and Technology & Artificial Intelligence, Nanjing Forestry University(信息科学与技术及人工智能学院,南京林业大学)
;
Department of Informatics, Modeling, Electronics, and Systems, University of Calabria(信息学、建模、电子与系统系,卡利博大学)
History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation
基于历史条件的时空视觉令牌剪枝用于高效视觉-语言导航
Qitong Wang, Yijun Liang, Ming Li, Tianyi Zhou, Christopher Rasmussen
机构
*
Department of Computer and Information Sciences at the University of Delaware(德克萨斯大学达勒姆分校计算机与信息科学系)
;
University of Maryland’s Department of Computer Science(马里兰大学计算机科学系)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation
MKG-RAG-Bench:多模态知识图谱增强生成中的检索基准
Xiaochen Wang, Bao Hoang, Han Liu, Ting Wang, Fenglong Ma
机构
*
The Pennsylvania State University(宾夕法尼亚州立大学)
;
Michigan State University(密歇根州立大学)
;
Dalian University of Technology(大连理工大学)
;
Stony Brook University(石溪大学)