机构
*
Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子及计算机工程学系)
;
Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程学系)
;
Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学医学智能与扩展现实研究所)
;
Center for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人创新中心)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构
*
Department of Computer Science, City University of Hong Kong (Dongguan)(香港城市大学(东莞)计算机科学系)
;
College of Computing and Data Science (CCDS), Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机科学学院)
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
Chan-Wei Hu, Yueqi Wang, Shuo Xing, Chia-Ju Chen, Suofei Feng, Ryan Rossi, Zhengzhong Tu
机构
*
Texas A&M University(德克萨斯农工大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
Stanford University(斯坦福大学)
;
Adobe Research(Adobe研究院)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
Hao Zhang, Chen Li, Basura Fernando
机构
*
Institute of High-Performance Computing, Agency for Science, Technology and Research, Singapore(高性能计算研究所,科学、技术及研究局,新加坡)
;
Centre for Frontier AI Research, Agency for Science, Technology and Research, Singapore(前沿人工智能研究中心,科学、技术及研究局,新加坡)
;
College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡)
CommentsCopyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
Multimodal LLM Integrated Semantic Communications for 6G Immersive Experiences
Yusong Zhang, Yuxuan Sun, Lei Guo, Wei Chen, Bo Ai, Deniz Gunduz
机构
*
School of Electronic and Information Engineering, Beijing Jiaotong University(电子信息工程学院,北京交通大学)
;
School of Electrical and Electronic Engineering, Imperial College London(电子电气工程学院,帝国理工学院伦敦分校)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG
CommentsThis work has been submitted to the IEEE for possible publication
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
Kim-Celine Kahl, Selen Erkan, Jeremias Traub, Carsten T. Lüth, Klaus Maier-Hein, Lena Maier-Hein, Paul F. Jaeger
机构
*
German Cancer Research Center (DKFZ) Heidelberg, Division of Medical Image Computing(德国癌症研究中心(DKFZ)海德堡,医学影像计算部门)
;
Helmholtz Imaging, German Cancer Research Center (DKFZ), Heidelberg, Germany(海德堡影像技术,德国癌症研究中心(DKFZ),海德堡,德国)
;
Faculty of Mathematics and Computer Science, University of Heidelberg, Germany(海德堡大学数学与计算机科学学院,德国)
;
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems(德国癌症研究中心(DKFZ)海德堡,智能医学系统部门)
;
German Cancer Research Center (DKFZ) Heidelberg, Interactive Machine Learning Group(德国癌症研究中心(DKFZ)海德堡,交互式机器学习小组)
;
Pattern Analysis and Learning Group, Department of Radiation Oncology, Heidelberg University Hospital, Germany(放射肿瘤科,海德堡大学医院模式分析与学习小组,德国)
;
National Center for Tumor Diseases (NCT) Heidelberg, Germany(海德堡国家肿瘤疾病中心(NCT),德国)