Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection
超越视觉取证:审计多模态鲁棒性用于合成医学图像检测
Ching-Hao Chiu, Hao-Wei Chung, Gelei Xu, Xueyang Li, Pin-Yu Chen, John Kheir, Meysam Ghaffari, Carlos Morato, Ahmed Abbasi, Yiyu Shi
机构
*
University of Notre Dame(圣母大学)
;
IBM Research(IBM研究院)
;
Boston Children’s Hospital(波士顿儿童医院)
;
Harvard Medical School(哈佛医学院)
;
Optum AI, UnitedHealth Group(Optum AI, 联合健康集团)
CommentsAccepted at MICCAI 2026. Version 2 is a substantial journal extension of the MICCAI 2026 conference version, with additional provenance perturbations, paired statistical analysis, extended SAVC mitigation experiments, and broader deployment discussion. 19 pages, 3 figures, 2 tables
机构
*
Information Technologies Institute, Centre for Research & Technology, Hellas(希腊研究与技术中心信息技术研究所)
;
Department of Electrical and Computer Engineering, Aristotle University of Thessaloniki(塞萨洛尼基亚里士多德大学电气与计算机工程系)
Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA
多模态大语言模型的置信度校准:基于医学视觉问答的实证研究
Yuetian Du, Yucheng Wang, Ming Kong, Tian Liang, Qiang Long, Bingdi Chen, Qiang Zhu
机构
*
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)
;
Zhihui Medical Technology (Shanghai) Co., Ltd.(智汇医疗科技(上海)有限公司)
机构
*
Built Environment Department, College of Science and Technology, North Carolina A&T State University(北卡罗来纳农工州立大学科技学院建筑环境系)
;
United Nations University Institute for Water, Environment and Health(联合国大学水、环境与健康研究所)
机构
*
Tsinghua University(清华大学)
;
Chongqing University(重庆大学)
;
Peking University(北京大学)
;
ZenoMind AI
;
Xi’an Jiaotong University(西安交通大学)
;
Beijing Institute of Technology(北京理工大学)
;
Southeast University(东南大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Joy Future Academy(京东探索研究院)
;
The University of Hong Kong(香港大学)
Frozen Multimodal Embeddings for AI-Assisted Interview Assessment of Personality and Cognitive Ability
冻结多模态嵌入用于异步视频面试中的个性与认知能力评估
Kuo-En Hung, Hung-Yue Suen, Shih-Ching Yeh, Hsiang-Wen Wang
机构
*
Technology Application and Human Resource Development, National Taiwan Normal University(台湾国立台中教育大学技术应用与人力资源发展系)
;
Computer Science and Information Engineering, National Central University(台湾国立中央大学计算机科学与资讯工程系)
;
Institute of Photonic System, National Yang Ming Chiao Tung University(台湾阳明交通大学光电系统研究所)
MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models
MMBU: 大规模多模态生物医学理解基准,用于探测视觉语言模型的感知能力
Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin, Min Woo Sun, Josiah Aklilu, James Burgess, Yuhui Zhang, Ryan Nayebi, Paola Avila, Robayo, Jin Ye, Ming Hu, Zhongying Deng, Junjun He, Xin Chen, Yue Yao, Robert Tibshirani, Jeffrey J. Nirschl, Serena Yeung-Levy
机构
*
Stanford University(斯坦福大学)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Instituto Tecnológico de Monterrey(蒙特雷技术学院)
;
Monash University(墨尔本大学)
;
University of Cambridge(剑桥大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Shandong University(山东大学)
机构
*
Beijing University of Technology(北京工业大学)
;
Central South University(中南大学)
;
Beijing Electronics Science & Technology Institute(北京电子科学技术研究所)
;
China Aerospace Science & Industry Corporation(中国航天科技集团)
;
Beijing Institute of Technology(北京理工大学)
;
Information Support Force Engineering University(信息支援部队工程大学)
;
Trunk Technology (Beijing) Co., Ltd.(trunk技术(北京)有限公司)
Comments18 pages, 4 figures, 2 tables. Published in the Proceedings of ASCAAD 2025
Journal refProceedings of the 13th International Conference of the Arab Society for Computation in Architecture, Art and Design (ASCAAD 2025), Riyadh, Saudi Arabia, 2025
Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning
对谁而言是可步行的?使用多模态深度学习捕捉步行感知中的主观变异性
Moloud Damandeh, Meead Saberi
机构
*
School of Civil and Environmental Engineering, University of New South Wales (UNSW)(新南威尔士大学土木与环境工程学院)
;
Research Centre for Integrated Transport Innovation (rCITI)(综合交通创新研究中心)
GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models
GraphVerse:面向多模态大语言模型的综合性视觉图推理基准
Yuanfu Sun, Yuanhang Ren, Kang Li, Chuanhao Ji, Jiaxi Li, Jiajin Liu, Ninghao Liu, Qiaoyu Tan
机构
*
New York University(纽约大学)
;
Sensetime Research(商汤科技研究院)
;
Tsinghua University(清华大学)
;
New York University Shanghai(上海纽约大学)
;
University of Georgia(佐治亚大学)
;
The Hong Kong Polytechnic University(香港理工大学)