Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
探索大语言模型编码器在胸部X光片图像-文本检索中的能力
Hanbin Ko, Gihun Cho, Inhyeok Baek, Donguk Kim, Joonbeom Koo, Changi Kim, Dongheon Lee, Chang Min Park
机构
*
Interdisciplinary Program in Bioengineering, Seoul National University Graduate School(生物工程跨学科项目,首尔国立大学研究生院)
;
Integrated Major in Innovative Medical Science, Seoul National University Graduate School(创新医学科学整合专业,首尔国立大学研究生院)
;
Department of Radiology, The First Affiliated Hospital, Zhejiang University School of Medicine(浙江大学医学院第一附属医院放射科)
;
Seoul National University College of Medicine(首尔国立大学医学院)
;
Department of Radiology, Seoul National University College of Medicine, Seoul National University Hospital(首尔国立大学医学院放射科,首尔国立大学医院)
;
Institute of Medical and Biological Engineering, Seoul National University Medical Research Center(医学与生物工程研究所,首尔国立大学医学研究所以及)
;
Institute of Radiation Medicine, Seoul National University Medical Research Center(放射医学研究所,首尔国立大学医学研究所以及)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
DenseMLLM:用于密集预测的标准多模态大语言模型
Yi Li, Hongze Shen, Lexiang Tang, Xin Li, Xinpeng Ding, Yinsong Liu, Deqiang Jiang, Xing Sun, Xiaomeng Li
机构
*
Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology, Hong Kong, China(香港科技大学电子与计算机工程系)
;
Tencent, Youtu-Lab, China(腾讯优图实验室)
Jailbreaking Multimodal Large Language Models using Multi-Clip Video
使用多片段视频破解多模态大语言模型
Choongwon Kang, Seungjong Sun, Hyunmin Jun, Jang Hyun Kim
机构
*
Department of Applied Artificial Intelligence, Sungkyunkwan University(应用人工智能系,成均馆大学)
;
Department of Human-Artificial Intelligence Interaction, Sungkyunkwan University(人机交互系,成均馆大学)
CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations
CardioLens: 通过多序列心脏MRI评估揭示MLLMs的临床现实差距
Zixian Su, Hongkai Zhang, Fan Gao, Encheng Su, Taiping Qu, Jingwei Guo, Nan Zhang, Hui Wang, Zhen Zhou, Kairui Bo, Yan Chen, Yue Ren, Shuai Li, Lei Xu, Henggui Zhang
机构
*
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Beijing Anzhen Hospital(北京安贞医院)
;
Beihang University(北航)
;
King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)
CommentsMICCAI 2026 Early Accept; Project Page: https://tahakoleilat.github.io/Evi-Steer. This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution will be published as part of the MICCAI 2026 proceedings in October
机构
*
China University of Petroleum (Beijing)(中国石油大学(北京))
;
Hainan Institute of China University of Petroleum (Beijing)(中国石油大学(北京)海南学院)
;
South China Normal University(华南师范大学)
Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy
超越文本与表格:ComProScanner中视觉-语言模型集成实现从科学图表中高精度提取材料数据
Aritra Roy, Enrico Grisan, Chiara Gattinoni, John Buckeridge
机构
*
Energy, Materials and Environment Research Centre, London South Bank University, London SE1 0AA, UK(能源、材料与环境研究中心,伦敦南银行大学)
;
School of Engineering and Design, London South Bank University, London SE1 0AA, UK(工程与设计学院,伦敦南银行大学)
;
Bioscience and Bioengineering Research Centre, London South Bank University, London SE1 0AA, UK(生物科学与生物工程研究中心,伦敦南银行大学)
;
Department of Physics, Kings College London, London WC2R 2LS, UK(物理系,伦敦国王学院)
CommentsProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) Accepted to ACL 2026 (Oral presentation). Code available at https://github.com/mkimhi/CARES
Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation
层级语义增强导航:面向视觉语言导航的最优传输与图驱动推理
Xiang Fang, Wanlong Fang, Changshuo Wang
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
;
Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore(新加坡南洋理工大学交叉学科研究生项目)
;
University College London(伦敦大学学院)
机构
*
Indian Institute of Technology Delhi, India(印度理工学院德里分校)
;
NVIDIA AI Technology Center, India(NVIDIA AI技术中心)
;
Jawaharlal Nehru University, India(贾瓦哈拉尔·尼赫鲁大学)
G2LoRA: Gradient Orthogonal Low-Rank Adaptation Framework for Graph Continual Learning on Text-Attributed Graphs
G2LoRA: 面向文本属性图的梯度正交低秩自适应框架用于图持续学习
Yuhan Wang, Yibo Ding, Yutong Ye, Mufan Zhao, Wenbo Zhang, Ruijie Wang, Jianxin Li
机构
*
School of Computer Science and Engineering, Beihang University(北航计算机科学与工程学院)
;
Department of Statistics, Columbia University(哥伦比亚大学统计系)
;
College of Computer Science, Beijing University of Technology(北京理工大学计算机学院)
机构
*
Tandon School of Engineering, New York University(纽约大学工程学院)
;
Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院)
;
Brookhaven National Laboratory(布鲁克海文国家实验室)