arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 1439 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 301 篇

2602.16144 2026-08-04 cs.CL cs.LG 版本更新 79%

Missing-by-Design: Certifiable Modality Deletion for Revocable Multimodal Sentiment Analysis

缺失-by-设计:可撤销多模态情感分析的可验证模态删除

Rong Fu, Ziming Wang, Chunlei Meng, Jiekai Wu, Kangan Qian, Hao Zhang, Simon Fong

机构 * University of Macau(澳门大学) Zhejiang University(浙江大学) Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) Juntendo University(立命馆大学) Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出MBD框架,通过结构化表示学习和可验证参数修改流程,实现可撤销多模态情感分析中的模态删除,实验表明其在不完整输入下具有强预测性能,并实现隐私与效用的平衡。

Comments 21 pages, 6 figures. In the previous version, Juntendo University was erroneously listed as the affiliation; we must clarify that this paper has absolutely no relation to Juntendo University. Therefore, we have replaced this affiliation in the new version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18668 2026-07-30 cs.LG cs.CV 版本更新 79%

On-Device Inference versus Wireless Streaming: Energy-Efficient Multi-Modal Deep Learning for Wearable Cardiovascular Patches

面向心血管传感器贴片的端到端多模态微型CNN原型设计

Mustafa Fuad Rifet Ibrahim, Tunc Alkanat, Felix Manthey, Maurice Meijer, Alexander Schlaefer, Peer Stelldinger

机构 * CTO System Innovation, NXP Semiconductors Germany GmbH(NXP半导体德国系统创新部) Advanced Chip Engineering, NXP Semiconductors(NXP半导体先进芯片工程部) Business Line Secure Connected Edge, NXP Semiconductors(NXP半导体安全连接边缘业务线) Institute of Medical Technology and Intelligent Systems, Hamburg University of Technology(汉堡技术大学医学技术与智能系统研究所) Department of Informatics, Hamburg University of Applied Sciences(汉堡应用科学大学信息学院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 针对资源受限的医疗边缘设备,提出一种早期融合心电图和心音图数据的卷积神经网络,实现二分类,相比现有技术将内存和计算成本降低约三个数量级,并验证了在微控制器上的能效优势。

Comments 16 pages, 2 figures. Extended version of our 2024 IEEE PerCom paper, with direct on-device energy measurements, a BLE communication benchmark, architecture comparisons, and an extended evaluation. Submitted to Pervasive and Mobile Computing; Measurement-method clarifications and minor editorial corrections; results and conclusions unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16129 2026-07-30 cs.AI 版本更新 79%

Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning

缩小眼科人工智能的差距:MM-Retinal-Reason数据集与OphthaReason模型用于动态多模态推理

Ruiqi Wu, Yuang Yao, Tengfei Ma, Chenran Zhang, Na Su, Tao Zhou, Geng Chen, Wen Fan, Yi Zhou

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Department of Ophthalmology, The First Affiliated Hospital of Nanjing Medical University(南京医科大学第一附属医院眼科学系) School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 该研究针对现有眼科AI仅聚焦基础推理的问题,构建首个眼科多模态数据集MM-Retinal-Reason,并提出带UADT方法的OphthaReason模型,实现眼科多模态推理性能的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02570 2026-07-28 eess.SP cs.AI 版本更新 79%

DOSE-I: A Multimodal Biosignal Dataset of Procedural Sedation for Endoscopy -- Technical Report

DOSE-I:用于内镜操作镇静的多模态生理信号数据集——技术报告

Jakob Garbe, Jan W. Kantelhardt, Katja Seeliger, Thomas Schmid

机构 * Universitätsmedizin Halle (Saale)(哈勒(萨勒)大学医学院) Universitätsklinikum Halle (Saale)(哈勒(萨勒)大学医院) Universitätsklinik und Poliklinik für Innere Medizin I(内科学I大学医院和多科医院) Martin‑Luther‑Universität Halle‑Wittenberg(哈勒-维滕贝格马尔伯特大学) Medizinische Fakultät(医学系) Ernst‑Grube‑Straße 40(埃尔斯特-格鲁布街40号)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文介绍内镜镇静多模态生理信号数据集DOSE-I的构建细节,提供标注数据、预处理方案与开源代码,支撑相关医疗AI研究。

Comments Dataset can be accessed via zenodo DOI https://doi.org/10.5281/zenodo.18483292 For citation use the primary academic reference: Garbe J et al. Towards predicting sedation depth in endoscopy with large clinically annotated EEG data of continuous Propofol sedation. In: P. Andreevetal (Eds.): AIME2026, LNAI 16749, p.1-6, Springer, 2026. https://doi.org/10.1007/978-3-032-30813-9_58

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15670 2026-07-27 cs.CV 版本更新 79%

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

PixDLM:一种用于无人机推理分割的双路径多模态语言模型

Shuyan Ke, Yifan Mei, Changli Wu, Yonghan Zheng, Jiayi Ji, Liujuan Cao, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出PixDLM,一种用于无人机推理分割的多模态语言模型,通过构建DRSeg基准数据集,验证了该模型在处理高分辨率无人机图像中的有效性。

Comments Accepted to CVPR 2026 (highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14675 2026-07-24 cs.RO cs.AI 版本更新 79%

An Intelligent-Cloud Edge Multimodal Interaction System for Robots

一种用于机器人的智能云边缘多模态交互系统

Zihan Guo, Xiaoqi Li

机构 * Hainan University(海南大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 针对复杂环境下资源受限机器人的交互问题,提出云边缘多模态交互框架,集成增强YOLO手势检测器与LLM、VLM智能体,改进手势检测方法,经实验验证该系统在手势检测精度、任务成功率及用户满意度方面表现良好,证明了方法的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14957 2026-07-21 cs.CV 版本更新 79%

Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging

学习用于多模态神经影像的稀疏潜在预测基础模型

Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian

机构 * New York University, Center for Data Science(纽约大学数据科学中心) NYU Grossman School of Medicine, Department of Radiology(纽约大学格罗斯曼医学院放射学系) State University of New York at Binghamton, School of Computing(纽约州立大学宾汉姆顿分校计算机学院) NYU Grossman School of Medicine, Department of Neurology(纽约大学格罗斯曼医学院神经病学系) NYU Grossman School of Medicine, Department of Neurosurgery(纽约大学格罗斯曼医学院神经外科学系) NYU Grossman School of Medicine, Department of Pathology(纽约大学格罗斯曼医学院病理学系) School of Medicine, Department of Radiology, Stanford(斯坦福大学医学院放射学系) NYU Grossman School of Medicine, Department of Neuroscience(纽约大学格罗斯曼医学院神经科学系) NYU Grossman School of Medicine, Neuroscience Institute(纽约大学格罗斯曼医学院神经科学研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 提出Neuro-JEPA模型,结合潜在预测目标和专家混合架构,学习T1w、T2w和FLAIR三种MRI序列的统一表示,在25项临床任务和22项公开数据集任务上优于现有基础模型和CNN基线。

Comments Under Review Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27958 2026-07-21 cs.AI 版本更新 79%

CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs

CARV:多模态大语言模型中组合类比推理的诊断基准

Yongkang Du, Xiaohan Zou, Minhao Cheng, Lu Lin

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出CARV任务及5500样本数据集,评估多模态大语言模型在组合类比推理中的能力,发现Gemini-2.5 Pro准确率仅为40.4%,远低于人类水平,揭示模型在规则提取与组合变换上的局限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10922 2026-07-21 cs.AI 版本更新 79%

Data-Efficient Curation for Multimodal Reasoning under Fixed Training Protocols

固定训练协议下多模态推理的数据高效整理

Yosub Shin, Michael Buriek, Boris Sobolev, Pavel Bushuyeu, Vikas Kumar, Haoyang Xu, Samuel Watson, Igor Molybog

机构 * University of Hawai'i at M\=anoa, Honolulu, HI, USA

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 研究固定协议微调下多模态推理的数据整理,以NeurIPS 2025 DCVLR挑战为测试平台,分析多种因素对推理准确性的影响,发现对齐源语料库的难度过滤增益最强,为数据受限的多模态推理微调提供经验方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10310 2026-07-20 cs.CL 版本更新 79%

PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment

PolyInterview:一个基于大语言模型的沉浸式模拟面试平台,具备全面的多模态评估

Zhiyuan Wen, Jiannong Cao, Kelly Chan, Zijian Wang, Chen Chen, Xiaoyun Liu, Jianing Yin, Zhuo Li

机构 * The Hong Kong Polytechnic University(香港理工大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 PolyInterview平台利用大语言模型,基于职位描述和简历为求职者生成定制面试问题,通过数字人类面试官进行多轮口语面试,全面评估回答内容、语音表达和非语言行为,提供结构化反馈,助力求职者更好地准备面试。

Comments 10 pages, 7 figures, and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04020 2026-07-20 cs.CV 版本更新 79%

Paired Uterine Whole-Slide Images and Pathology Reports for Multimodal Computational Pathology

用于多模态计算病理学的配对子宫全切片图像和病理报告

Han Li, Jingsong Liu, Ayako Ura, Junlin Hou, Zhengyang Xu, Azar Kazemi, Oskar Thaeter, Christian Grashei, Fabian Gülhan, Reza Nasirigerdeh, Xun Ma, Rui Yan, Hao Chen, S. Kevin Zhou, Nassir Navab, Carolin Mogler, Peter Schüffler

机构 * Institute of Pathology, Technical University of Munich(慕尼黑工业大学病理研究所) Computer Aided Medical Procedures (CAMP), Technical University of Munich(慕尼黑工业大学计算机辅助医疗程序(CAMP)) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Department of Human Pathology, Juntendo University Graduate School of Medicine(顺天堂大学医学研究生院人体病理学部) The Hong Kong University of Science and Technology(香港科技大学) Munich Data Science Institute (MDSI)(慕尼黑数据科学研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 研究子宫疾病病理诊断,针对全切片图像与病理报告配对数据集稀缺问题,引入TUM-Uteria数据集,含多对病例及切片级配对,经验证,为计算病理学研究提供支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13800 2026-07-17 cs.CV 版本更新 79%

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

超越医学诊断:医学多模态大语言模型如何在空间中思考

Quoc-Huy Trinh, Xi Ding, Yang Liu, Zhenyue Qin, Xingjian Li, Gorkem Durak, Halil Ertugrul Aktas, Andrea M. Bejar, Ulas Bagci, Min Xu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出SpatialMed基准,通过自主合成空间视觉问答数据评估医学MLLMs的3D空间智能,发现现有模型在医学影像空间推理能力不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06285 2026-07-09 cs.CV 版本更新 79%

MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training

MMEarth-Bench:通过多模态测试时训练进行全局模型适配

Lucia Gordon, Serge Belongie, Christian Igel, Nico Lang

机构 * Harvard University, USA(哈佛大学,美国) University of Copenhagen, Denmark(哥本哈根大学,丹麦)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 研究针对地理空间机器学习中现有基准数据集不足,引入含多模态任务的MMEarth-Bench,通过多模态测试时训练方法提升模型性能,改善地理泛化能力。

Comments Published at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04628 2026-07-08 cs.CV 版本更新 79%

A Spatial-Spectral-Frequency Interactive Network for Multimodal Remote Sensing Classification

一种空间-光谱-频率交互网络用于多模态遥感分类

Hao Liu, Yunhao Gao, Wei Li, Mingyang Zhang, Maoguo Gong, Lorenzo Bruzzone

机构 * Department of Information Engineering and Computer Science, University of Trento(特伦托大学信息工程与计算机科学系) School of Information and Electronics, Beijing Institute of Technology(北京理工大学信息与电子学院) Beijing Key Laboratory of Fractional Signals and Systems, Beijing Institute of Technology(北京理工大学分数域信号与系统北京市重点实验室) School of Electronic Engineering, Xidian University(西安电子科技大学电子工程学院) Key Laboratory of Collaborative Intelligent Systems of Ministry of Education, Xidian University(西安电子科技大学教育部协同智能系统重点实验室) Academy of Artificial Intelligence, Inner Mongolia Normal University(内蒙古师范大学人工智能研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出S$^2$Fin网络,通过空间、光谱和频率域的交互模块,提升多模态遥感图像分类性能,实验表明其在有限标注数据下表现优于现有方法。

Journal ref Pattern Recognition, Vol. 180, 2026, 114309

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09696 2026-07-08 cs.CV 版本更新 79%

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

ZeroBench:当代大型多模态模型的一个不可能的视觉基准测试

Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma, Akash Gupta, Samuel Roberts, Ioana Croitoru, Simion-Vlad Bogolin, Jialu Tang, Florian Langer, Vyas Raina, Vatsal Raina, Hanyi Xiong, Vishaal Udandarao, Jingyi Lu, Shiyang Chen, Sam Purkis, Tianshuo Yan, Wenye Lin, Gyungin Shin, Qiaochu Yang, Anh Totti Nguyen, David I. Atkinson, Aaditya Baranwal, Alexandru Coca, Mikah Dang, Sebastian Dziadzio, Jakob D. Kunz, Kaiqu Liang, Alexander Lo, Brian Pulfer, Steven Walton, Charig Yang, Kai Han, Samuel Albanie

机构 * University of Cambridge(剑桥大学) University of Alberta(阿尔伯塔大学) The University of Hong Kong(香港大学) University of Oxford(牛津大学) Northeastern University(东北大学) Astadeus College of Southern Maryland(马里兰州南部学院) University of Geneva(日内瓦大学) University of Oregon(俄勒冈大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 针对大型多模态模型在图像解释方面不足但在现有视觉基准测试中优势易逝的问题,引入通过对抗过滤策划的ZeroBench基准测试,评估多个模型,展现其潜力并公开测试内容,为模型视觉能力评估提供长期有效的新基准。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30673 2026-07-07 cs.CL 版本更新 79%

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

TeachObs:多模态教学观察与模型评估的人工验证基准

Yeil Jeong, Youngjin Yoo, Jiyoung Bae, Seobin Sohn, Hyejin Han, Jinseo Lee, Howard Scott, Unggi Lee

机构 * Indiana University Bloomington(印第安纳大学布卢明顿分校) Pai Chai University(培才大学) Seoul National University(首尔国立大学) Ewha Womans University(成均馆大学) University of Wolverhampton(沃尔夫汉普顿大学) Korea University Sejong Campus(韩国大学世宗校区)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 提出TeachObs基准,包含30节公开课视频的5158个15秒场景,由7名研究者标注39个二值观察码,并评估5个前沿视觉大语言模型在三种任务上的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22867 2026-07-07 cs.CV cs.RO 版本更新 79%

MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments

MUSON:用于城市环境中社会合规导航的面向推理的多模态数据集

Zhuonan Liu, Xinyu Zhang, Zishuo Wang, Runji Cai, Tomohito Kawabata, Qianyi Li, Xuance Peng, Tianze Yu, Zhen Xiong, Xuesu Xiao, Ling Xiao

机构 * Graduate School of Information Science and Technology, Hokkaido University(北海道大学信息科学研究生院) Graduate School of Computer Science, George Mason University(乔治·马歇尔大学计算机科学研究生院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 研究社会合规导航问题,利用视觉语言模型,引入含10110个自我中心样本的多模态数据集MUSON,采用五步注释框架,为推进该导航提供有效基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20272 2026-07-07 cs.CV 版本更新 79%

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

VKnowU:评估多模态语言模型中的视觉知识理解

Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu, Xiangyu Zeng, Limin Wang, Yu Qiao, Yi Wang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai AI Laboratory(上海人工智能实验室) Nanjing University(南京大学) Shanghai Innovation Institute(上海创新研究院) City University of Hong Kong(香港城市大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 研究多模态大语言模型视觉知识理解能力,提出VKnowU基准测试。评估28个模型,发现与人类表现有差距。引入新数据集和VideoKnow+基线模型,采用结构化范式和强化学习,提升模型表现。

Comments Accepted by ECCV 2026. https://github.com/OpenGVLab/VKnowU

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15379 2026-07-07 cs.CV 版本更新 79%

The P$^3$ Dataset: Pixels, Points and Polygons for Multimodal Building Vectorization

P$^3$数据集:用于多模态建筑物矢量化的像素、点和多边形

Raphael Sulzer, Liuyun Duan, Nicolas Girard, Florent Lafarge

机构 * Université Côte d’Azur, INRIA(法国里沃利大学、INRIA) LuxCarta Technology(LuxCarta技术公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 介绍用于建筑物矢量化的多模态P$^3$数据集,由多源数据构建,含大量高精度激光雷达点等。证明激光雷达点云可用于预测建筑物多边形,融合激光雷达和图像能提高预测精度与几何质量,数据集公开。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17360 2026-07-03 cs.CV 版本更新 79%

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

Omni-DuplexEval: 评估实时双工全模交互

Chaoqun He, Mingyang Xiang, Yingjing Xu, Bokai Xu, Junbo Cui, Jie Zhou, Yuan Yao, Lijie Wen

机构 * Tsinghua University(清华大学) Tongji University(同济大学) ModelBest Inc.(ModelBest公司)

专题命中 多模态评测 :omni-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出Omni-DuplexEval基准,用于系统评估实时双工交互能力,通过两个互补场景评估模型生成连续响应和主动提醒的能力,并揭示现有模型在平衡响应及时性和内容连贯性方面的局限性。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19614 2026-07-02 cs.LG cs.CV 版本更新 79%

Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

多重性是多模态学习中不可避免且固有的挑战

Sanghyuk Chun, Olga Russakovsky

机构 * Princeton University(普林斯顿大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本文论证多模态数据固有的多对多映射关系(多重性)是影响数据构建、模型训练和评估的根本瓶颈,并呼吁发展多重性感知的学习框架与评估协议。

Comments ICML 2026 Position Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19466 2026-06-29 cs.CV 版本更新 79%

ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models

ProactiveBench: 多模态大语言模型主动性的基准测试

Thomas De Min, Subhankar Roy, Stéphane Lathuilière, Elisa Ricci, Massimiliano Mancini

机构 * University of Trento(特伦托大学) University of Bergamo(贝拉米奥大学) Inria Grenoble(格勒诺布尔研究所) Bruno Kessler Foundation(布鲁诺·凯斯勒基金会)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 提出ProactiveBench基准,评估多模态大语言模型在遮挡识别等任务中请求用户干预的主动性,发现模型普遍缺乏主动性且与能力无关,但可通过强化学习微调习得。

Comments Accepted at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14837 2026-06-23 cs.CV 版本更新 79%

DamageArbiter: A Multimodal Arbitration Framework for Disaster Damage Assessment from Street-View Imagery

DamageArbiter:一种基于街景图像进行灾害损伤评估的多模态仲裁框架

Yifan Yang, Lei Zou, Wenjing Gong, Kani Fu, Zongrong Li, Siqin Wang, Bing Zhou, Heng Cai, Hao Tian

机构 * organization= Department of Geography, Texas A\&M University , city= College Station , country= USA organization= Department of Landscape Architecture \& Urban Planning, Texas A\&M University , city= College Station , country= USA organization= Department of Industrial Systems Engineering, University of Florida , city= Gainesville , country= USA organization= Spatial Sciences Institute, University of Southern California , city= Los Angeles , country= USA organization= Department of Geography Sustainability, University of Tennessee , city= Knoxville , country= USA

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 提出DamageArbiter多模态仲裁框架,通过轻量级逻辑回归元分类器仲裁单模态与多模态模型预测分歧,在2556张街景图像上将准确率提升至75.85%,MCC达0.6188,并将过度自信误差从70.58%降至16.45%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28387 2026-06-19 cs.AI cs.LG 版本更新 79%

The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation

脚手架效应:提示框架如何驱动临床VLM评估中的表面多模态增益

Doan Nam Long Vu, Simone Balloccu

机构 * Technical University of Darmstadt(达姆施塔特技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 研究发现,在临床VLM评估中,提示中提及MRI可用性即可解释70-80%的性能提升,与图像数据是否存在无关,这种“脚手架效应”揭示了表面评估无法反映真实多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05368 2026-06-18 cs.CV 版本更新 79%

Biomazon: A Multimodal Dataset for 3D Forest Structure and Biomass Modeling in the Amazon Basin

Biomazon:亚马逊盆地三维森林结构与生物量建模的多模态数据集

Sayan Mandal, Rocco Sedona, Simon Besnard, Mikhail Urbazaev, Morris Riedel, Ehsan Zandi, Gabriele Cavallaro

机构 * Jülich Supercomputing Centre (JSC), Forschungszentrum Jülich(julich超级计算中心(JSC),julich研究所) School of Engineering and Natural Sciences (SENS), University of Iceland(工程与自然科学学院(SENS),冰岛大学) Global Land Monitoring Group, GFZ Helmholtz Centre for Geosciences(全球土地监测组,geofz赫尔姆霍兹研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 针对现有方法未将森林垂直结构作为有序轮廓学习的问题,提出Biomazon多模态基准数据集,结合GEDI RH和AGBD目标与多传感器预测因子,通过共享编码器-解码器框架进行消融研究,为热带森林结构一致RH轮廓预测和结构-生物量建模建立参考基准。

Comments 32 pages, 21 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18109 2026-06-18 cs.CL cs.SD 版本更新 79%

FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings

FLiP:理解和解释多模态多语句子嵌入

Santosh Kesiraju, Bolaji Yusuf, Šimon Sedláček, Oldřich Plchot, Petr Schwarz

机构 * Brno University of Technology(布拉格技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 提出因子化线性投影(FLiP)模型,从多语言、多模态句子嵌入中恢复词汇内容,揭示编码器的模态和语言偏差。

Comments Accepted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10384 2026-06-17 cs.CL 版本更新 79%

When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents

当表格失控:评估多模态模型在法语金融文档上的表现

Virginie Mouilleron, Théo Lasnier, Anna Mosolova, Djamé Seddah

机构 * Inria Paris(巴黎国家信息与自动化研究所) Sorbonne Université(索邦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 提出Scribe Finance基准,评估多模态模型在法语金融文档上的文本、表格、图表及多轮对话理解能力,发现模型在图表和多轮对话中表现脆弱。

Comments 16 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18827 2026-06-16 q-bio.NC cs.AI 版本更新 79%

OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens

OmniMouse: 基于1500亿神经令牌的多模态多任务脑模型的可扩展性

Konstantin F. Willeke, Polina Turishcheva, Alex Gilbert, Goirik Chakrabarty, Hasan A. Bedel, Paul G. Fahey, Yongrong Qiu, Marissa A. Weis, Michaela Vystrčilová, Taliah Muhammad, Lydia Ntanavara, Rachel E. Froebe, Kayla Ponder, Zheng Huan Tan, Emin Orhan, Erick Cobos, Sophia Sanborn, Katrin Franke, Fabian H. Sinz, Alexander S. Ecker, Andreas S. Tolias

机构 * Department of Ophthalmology, Byers Eye Institute, Stanford University(斯坦福大学眼科学系、比尔斯眼科研究所) Stanford Bio-X, Stanford University(斯坦福大学生物交叉学科) Wu Tsai Neurosciences Institute, Stanford University(斯坦福大学吴泰教授神经科学研究所) Institute of Computer Science and Campus Institute Data Science, University Göttingen(哥廷根大学计算机科学研究所和校园数据科学研究所)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

AI总结 利用小鼠视觉皮层31亿神经元数据,训练多模态多任务模型OmniMouse,在神经预测、行为解码等任务上达到最优,发现性能随数据量可靠提升但模型规模收益饱和,与AI领域标准扩展规律相反。

Comments Published at ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22391 2026-06-16 cs.CL 版本更新 79%

Detecting Hate and Inflammatory Content in Bengali Memes: A New Multimodal Dataset and Co-Attention Framework

检测孟加拉语模因中的仇恨和煽动性内容:一个新的多模态数据集和共注意力框架

Rakib Ullah, Mominul islam, Md Sanjid Hossain, Md Ismail Hossain

机构 * University of California, Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CL

AI总结 针对孟加拉语模因中仇恨和煽动性内容检测的研究空白,构建了首个区分煽动性内容与直接仇恨言论的数据集Bn-HIB,并提出多模态共注意力融合模型MCFM,通过联合分析视觉和文本特征实现更准确分类。

Comments Added public link to dataset and fixed typo in abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05888 2026-06-16 cs.CV 版本更新 79%

BioAutoML-NAS: An End-to-End AutoML Framework for Multimodal Insect Classification via Neural Architecture Search on Large-Scale Biodiversity Data

BioAutoML-NAS:基于大规模生物多样性数据通过神经架构搜索进行多模态昆虫分类的端到端AutoML框架

Arefin Ittesafun Abian, Debopom Sutradhar, Md Rafi Ur Rashid, Reem E. Mohamed, Md Rafiqul Islam, Asif Karim, Kheng Cher Yeo, Sami Azam

机构 * Department of Computer Science and Engineering, United International University, Dhaka, Bangladesh(乌姆国际大学计算机科学与工程系,达卡,孟加拉国) Applied Artificial Intelligence and Intelligent Systems (AAIINS) Laboratory, 1217, Dhaka, Bangladesh(应用人工智能与智能系统实验室(AAIINS),1217号,达卡,孟加拉国) Department of Computer Science and Engineering, Penn State University, University Park, PA, USA(宾夕法尼亚州立大学计算机科学与工程系,University Park,PA,美国) Faculty of Science and Information Technology, Charles Darwin University, Sydney, NSW, Australia(查尔斯达尔文大学科学与信息技术学院,悉尼,新南威尔士州,澳大利亚) Faculty of Science and Technology, Charles Darwin University, Casuarina, 0909, NT, Australia(查尔斯达尔文大学科学与技术学院,Casuarina,0909,北领地,澳大利亚)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 提出首个多模态BioAutoML模型BioAutoML-NAS,利用神经架构搜索自动学习图像操作,结合元数据融合与交替双层优化,在BIOSCAN-5M数据集上以96.81%准确率超越现有方法。

Comments Accepted in IEEE Transactions on Big Data

Journal ref IEEE Transactions on Big Data (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏