arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-07-31 至 2026-07-31 共收录 17 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 17 篇

2607.27667 2026-07-31 cs.CV 新提交 87%

Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers

证据见证组合:针对闭集多模态答案的单预填充风险检测

Fexiang Liu, Shiye Wang, Qiang Qiu, Zheng Wang

专题命中 多模态评测 :multimodal(title,abstract);MLLM(summary_cn,abstract_cn);分类 cs.CV

AI总结 该研究提出WEP方法,利用MLLM的白盒预填充路径,无需额外操作即可检测闭集视觉答案的推理风险,在3个MLLM和4个基准上提升了平均错误AP。

Comments 22 pages, 6 figures; includes supplementary material. Code: https://github.com/SouthWinter/WEP

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28318 2026-07-31 cs.AI 新提交 85%

PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

PathView-Bench:多模态大语言模型能否实现病理图像的细粒度多尺度理解?

Zongyi Chen, Yu Liang, Jie Lin, Liansheng Wang

机构 * National Institute for Data Science in Health and Medicine, Xiamen University(厦门大学健康与医学数据科学国家研究院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.AI

AI总结 研究针对现有病理多模态基准的不足,推出PathVU基准,评估发现主流MLLMs在多尺度病理图像细粒度视觉任务上存在显著局限,为相关模型的开发评估提供了可复现基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28269 2026-07-31 cs.CV cs.AI cs.MM 新提交 85%

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation

Theia:用于无数据蒸馏的Incidents1M数据集的大规模多模态字幕生成与自动验证

Simone Giano, Lorenzo Severini, Alessandro Galdelli, Adriano Mancini

机构 * Università Politecnica delle Marche(马尔凯理工大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本研究针对灾害领域多模态数据集的缺陷,提出方法构建并自动验证Incidents1M数据集,生成高保真字幕,其语义一致性达78.65/100,为跨模态知识蒸馏提供了可扩展框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25991 2026-07-31 cs.AI cs.CV 版本更新 84%

Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline

面向社交媒体中统一的多模态虚假信息检测:基准数据集与基线模型

Haiyang Li, Yaxiong Wang, Shengeng Tang, Yuchen Zhang, Lianwei Wu, Lechao Cheng, Liu Liu, Chaofeng Dong, Zhun Zhong

机构 * School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学) School of Computer Science and Technology, Northwestern Polytechnical University(计算机科学与技术学院,西北工业大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 该研究构建了含9.8万样本的OmniFake基准数据集,提出UMFDet框架,实现对人工与AI生成两类多模态虚假内容的统一检测,性能优于专用基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04281 2026-07-31 cs.AI 版本更新 83%

RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model

RetiBridge:用知识引导的多模态大语言模型连接定量视网膜生物标志物与定性诊断

Zhuangzhi Gao, Hongyi Qin, He Zhao, Qinkai Yu, Feixiang Zhou, Fu Wang, Jinru Ding, Eduard Shantsila, Uazman Alam, Alena Shantsila, Wahbi El-Bouri, Gregory Y. H. Lip, Yalin Zheng

机构 * University of Liverpool(利物浦大学) Institute of Life Course & Medical Sciences(生命课程与医学科学研究院) Department of Eye and Vision Sciences(眼科与视觉科学系) Computer Science Department(计算机科学系) Cardiovascular & Metabolic Medicine(心血管与代谢医学) Liverpool Centre for Cardiovascular Science(利物浦心血管科学中心)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract_cn);分类 cs.AI

AI总结 该研究提出RetiBridge模型,结合CFP、OCT数据与文本,通过知识引导微调学习定量到定性诊断路径,在眼科理解基准上表现优于开源基线及OpenAI o3,代码数据已公开。

Comments 10 pages, 4 figures, 3 table. Equal contribution: Zhuangzhi Gao and Hongyi Qin. Corresponding author: Yalin Zheng (yzheng@liverpool.ac.uk)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27378 2026-07-31 cs.CV cs.MM 新提交 82%

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

PanDent:面向牙科放射学中全面的牙级结构-语言一致性

Xiaohan Li, Xinyu Liu, Chang Liu, Sum Wing Au Yeung, Jun Liu, Yixuan Yuan, Hui Chen

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙医学院) Imperial College London(帝国理工学院) University of Science and Technology of China(中国科学技术大学) Department of Data and Systems Engineering, The University of Hong Kong(香港大学数据与系统工程系) Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系)

专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.MM

AI总结 本研究推出PanDent牙科OPG基准,经实验发现现有MLLM生成的牙科报告流畅但临床一致性差,在PanDent上微调可提升其结构-语言一致性,该基准可用于评估MLLM的牙级临床推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27393 2026-07-31 cs.CL cs.AI 新提交 81%

AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

AHA-Memes:用于理解阿拉伯语表情包中仇恨内容的细粒度多模态基准

Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Abul Hasnat, Md. Rafiul Biswas, Wajdi Zaghouani, Firoj Alam

机构 * Qatar Computing Research Institute(卡塔尔计算研究所) Hamad Bin Khalifa University(哈马德·本·哈利法大学) Northwestern University in Qatar(卡塔尔西北大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 该研究推出首个带细粒度多标签标注的阿拉伯语仇恨表情包基准AHA-Memes,构建含5000张人工标注及6.6万张银标的数据集,对多类模型基准测试并发布资源以推动相关研究。

Comments 26 pages, 14 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27069 2026-07-31 cs.CV cs.AI 版本更新 81%

Visual Credit Audit for Multimodal Spatial Reasoning

多模态空间推理的视觉信用审计

Feixiang Liu, Qiang Qiu, Lanbo Sun, Nan Wei, Huawei Shen, Xueqi Cheng

专题命中 多模态评测 :multimodal(title);MLLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 该研究提出视觉信用审计(VCA)方法,分解多模态空间推理基准的成功维度,发现部分模型决策正确但未获图像额外信用,验证了其在评估模型视觉关系响应上的有效性。

Comments 20 pages, 2 figures. Code: https://github.com/SouthWinter/VCA

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27357 2026-07-31 cs.CV eess.IV 新提交 79%

Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification

用于非配对跨模态医学分类的共享语义码本蒸馏

Dillan Imans, Phuoc-Nguyen Bui, Duc-Tai Le, Hyunseung Choo

机构 * Sungkyunkwan University(成均馆大学) Convergence Research Institute, Sungkyunkwan University(成均馆大学融合研究院)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

AI总结 针对非配对跨模态医学分类的挑战,提出SSCD方法,通过共享语义码本转移知识,在两个医学分类任务中提升学生模型性能,优于各基线方法

Comments 16 pages, 2 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27296 2026-07-31 cs.SD cs.MM 新提交 79%

SKY-Piano: A Multimodal Piano Performance Dataset

SKY-Piano:多模态钢琴演奏数据集

Joonhyung Bae, Dawon Park, Taegyun Kwon, Yoon-Seok Choi, Hyeon Hur, Satoshi Obata, Shigeru Kai, Yohei Wada, Yu Takahashi, Akira Maezawa, Jaebum Park, Jonghwa Park, Juhan Nam

机构 * Seoul National University(首尔大学) KAIST(韩国科学技术院) Yamaha(雅马哈)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.MM

AI总结 该研究发布SKY-Piano多模态钢琴演奏数据集,含多类数据及标注工具,通过微调实验验证其用于MIDI到动作生成的可用性。

Comments Accepted to the 27th International Society for Music Information Retrieval Conference (ISMIR 2026), Abu Dhabi, UAE. Project page: https://joonhyungbae.github.io/skypiano/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27856 2026-07-31 cs.CV 新提交 70%

Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation

基准测试用于小样本医学图像分割的基础模型与大语言模型

Jinghong Liu, Yuchuan Deng, Fanping Liu, Meng Huang, Xirong Li

专题命中 多模态评测 :MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 该研究推出统一基准FAME,涵盖四类模型,含14958个样本,评估得出小样本分割的关键规律,助力开发更有效的小样本医学分割方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27895 2026-07-31 cs.AI cs.CV 新提交 62%

MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

MMHBench:用于长视频心理健康理解的多视角基准测试集

Jinpeng Hu, Erqiang Wang, Shan Wang, Zhuo Li, Peipei Song, Xun Yang, Meng Wang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出MMHBench多视角心理健康理解基准,含268个长视频与2184条问题,采用MAQG框架生成问题,评估22个多模态大语言模型后发现该任务仍具挑战性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28401 2026-07-31 cs.CV 新提交 57%

Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis

面向操作制图与影响分析的大规模跨区域遥感洪水监测框架

Ilya Novikov, Svetlana Illarionova, Ruslan Dzharkinov, Maria Smirnova, Ayrat Abdullin, Anna Korotkova, Mariia Ulianova, Dmitrii Shadrin, Evgeny Burnaev

机构 * Skolkovo Institute of Science and Technology(斯科尔科沃科学技术研究所) Trofimuk Institute of Petroleum Geology and Geophysics SB RAS(俄罗斯科学院西伯利亚分院特罗菲穆克石油天然气地质与地球物理研究所) King Fahd University of Petroleum and Minerals(法赫德国王石油与矿产大学) Tyumen Industrial University(秋明工业大学) Huawei Russian Research Institute(华为俄罗斯研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本研究提出端到端多模态洪水监测框架,对比U-Net++与AnySat两种水面检测策略,结合多模态卫星数据估算洪水影响,在2019年图伦洪水监测中表现良好,为跨区域大规模洪水监测提供可行方案。

Comments 37 pages, 11 figures, 9 tables. Preprint submitted to Earth Systems and Environment. This version has not been peer reviewed

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28187 2026-07-31 cs.AI cs.CR cs.SE 新提交 57%

Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation

旧技巧,新模型:简单图像变换如何破解基于现代AI的内容审核

Marco Alecci, Francesco Marchiori, Iyiola Emmanuel Olatunji, Tegawendé F. Bissyandé, Jacques Klein

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 该研究通过黑盒评估发现,商用多模态图像审核API可被简单图像变换绕过,且不同服务稳健性差异大,表明这类API无法单独作为可靠安全过滤器,需结合分层审核管道部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27806 2026-07-31 cs.CV 新提交 57%

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

LoMeVQA:纵向医学视觉问答综合基准

Zhilin Wu, Zhangkai Ni, Chengmei Yang, Longzhen Yang, Yihang Liu, Ying Wen, Lianghua He

机构 * Tongji University(同济大学) East China Normal University(华东师范大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 该研究提出纵向医学视觉问答基准LoMeVQA,发现现有多模态大语言模型在该任务上时间推理能力不足,推出MedLong-8B实现最优性能,并开展相关分析。

Comments 23 pages, 17 figures, 7 tables. Code and data: https://github.com/pepperbubble/LoMeVQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00448 2026-07-31 cs.CV eess.IV 版本更新 57%

Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis

从压缩CT学习:用于资源高效医学图像分析的特征注意力风格迁移和结构化因子化投影

Shadid Yousuf, S. M. Mahbubur Rahman, Mohammed Imamul Hassan Bhuiyan

机构 * Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology(电气电子工程系,孟加拉工程与技术大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出FAST和SFP方法,通过压缩CT体积进行胸腔异常检测,实现低资源部署和高效数据传输,实验表明CT-Lite在压缩输入下达到接近未压缩基线的AUROC性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27217 2026-07-31 stat.AP cs.LG 新提交 50%

Foundation-Model Earth Representations Enable Regional-Scale Forest Aboveground Biomass Monitoring Across the Northeastern United States

基于基础模型的地球表征实现美国东北部区域尺度森林地上生物量监测

Shashika Lamahewage, Chandi Witharana

专题命中 多模态评测 :multimodal(abstract)

AI总结 该研究利用AlphaEarth基础模型生成的Google卫星嵌入,结合LiDAR数据与森林清查数据构建模型,实现美国东北部区域森林地上生物量的高精度监测,为规模化碳评估提供了新途径。

详情

展开后加载摘要…

URL PDF HTML 收藏