arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-07-09 至 2026-07-09 共收录 17 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 17 篇

2607.01420 2026-07-09 cs.CL cs.AI cs.CV 新提交 82%

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

MultAttnAttrib:长文档问答中的无训练多模态归因

Dang Quang Thien Tran, Quang V. Dang, Vinamra Tyagi, Sai Soorya Rao Veeravalli, Trang Nguyen, Ryan A. Rossi, Franck Dernoncourt, Nedim Lipka, Koustava Goswami, Samyadeep Basu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出无训练归因方法MultAttnAttrib,利用模型预填充、选定注意力头和校准阈值定位证据,并构建多模态归因基准MultAttrEval,实验表明其优于多种方法,匹配GPT 5.4且延迟降低至1/7。

Comments 25 pages (8 main, 17 references + appendix), 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06633 2026-07-09 cs.CV cs.AI 新提交 81%

ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities

ProMoE-FL:用于具有缺失模态的多模态联邦学习的原型条件专家混合模型

Aavash Chhetri, Bibek Niroula, Eduard Vazquez, Yash Raj Shrestha, Prashnna Gyawali, Loris Bazzani, Binod Bhattarai

机构 * NepAl Applied Mathematics and Informatics Institute(尼泊尔应用数学与信息学研究所) Fogsphere (Redev.AI Ltd)(福格球(Redev.AI有限公司)) University of Lausanne(洛桑大学) West Virginia University(西弗吉尼亚大学) University of Verona(维罗纳大学) University College London(伦敦大学学院) University of Aberdeen(阿伯丁大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 针对多模态联邦学习中缺失模态问题,提出ProMoE-FL框架,构建全局客户端感知原型库,以原型和模态索引为条件实现专家路由来动态合成缺失特征,在四个胸部X光数据集评估中优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03740 2026-07-09 q-bio.QM cs.CV eess.IV 新提交 79%

Triple-Phase Multimodal Knowledge Aggregation Framework for Microbial Keratitis Subtype Diagnosis on Slit-Lamp Photography

基于裂隙灯摄影的微生物性角膜炎亚型诊断三相多模态知识聚合框架

Yiqing Wang, Maria A. Woodward, Ziyun Yang, N. Venkatesh Prajna, Chunming He, Leslie M. Niziol, Mercy Pawar, Ming-Chen Lu, Guillermo Amescua, Rachel Wozniak, Sejal Amin, Abinaya Krishnan, Prabhleen Kochar, Sina Farsiu

机构 * Department of Biomedical Engineering, Duke University(杜克大学生物医学工程系) Kellogg Eye Center, Department of Ophthalmology and Visual Sciences, University of Michigan(密歇根大学凯洛格眼科中心,眼科学与视觉科学系) Department of Cornea and Refractive Surgery Services, Aravind Eye Care System(阿瓦因眼科医疗系统角膜与屈光手术部) Bascom Palmer Eye Institute, Department of Ophthalmology, University of Miami Miller School of Medicine(迈阿密大学米勒医学院巴斯科姆·帕勒眼科研究所,眼科学系) Flaum Eye Institute, Department of Ophthalmology, University of Rochester Medical Center(罗切斯特大学医学中心弗劳姆眼科研究所,眼科学系) Department of Ophthalmology, Henry Ford Hospital(亨利福特医院眼科部) Duke Eye Center, Duke University School of Medicine(杜克大学医学院杜克眼科中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 针对传统微生物性角膜炎病原体鉴定慢、成本高的问题,提出融合多模态信息的三相框架实现角膜炎亚型分类,经多中心验证性能优异,跨站点泛化评估更具参考性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06285 2026-07-09 cs.CV 版本更新 79%

MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training

MMEarth-Bench:通过多模态测试时训练进行全局模型适配

Lucia Gordon, Serge Belongie, Christian Igel, Nico Lang

机构 * Harvard University, USA(哈佛大学,美国) University of Copenhagen, Denmark(哥本哈根大学,丹麦)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 研究针对地理空间机器学习中现有基准数据集不足,引入含多模态任务的MMEarth-Bench,通过多模态测试时训练方法提升模型性能,改善地理泛化能力。

Comments Published at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06973 2026-07-09 cs.LG 新提交 78%

Rethinking Multimodal Time-Series Forecasting Evaluation

重新思考多模态时间序列预测评估

Haoxin Liu, Yichen Zhou, Rajat Sen, B. Aditya Prakash, Abhimanyu Das

机构 * Georgia Institute of Technology(佐治亚理工学院) Google Research(谷歌研究院)

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 研究针对现有多模态时间序列预测基准的问题,引入TimesX基准,通过自动数据生成管道获取多样真实世界时间序列。经实证研究发现,一些在现有基准上好的方法在TimesX上可能失败,而利用文本上下文的简单集成方法表现出色。

Journal ref Published in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18164 2026-07-09 cs.RO 版本更新 78%

GrandTour: A Legged Robotics Dataset in the Wild for Multi-Modal Perception and State Estimation

GrandTour: 一种用于多模态感知与状态估计的野外腿式机器人数据集

Turcan Tuna, Jonas Frey, Frank Fu, Katharine Patterson, Tianao Xu, Maurice Fallon, Cesar Cadena, Marco Hutter

机构 * ETH Zurich(苏黎世联邦理工学院) Stanford University(斯坦福大学) UC Berkeley(伯克利大学) Oxford University(牛津大学)

专题命中 多模态评测 :multi-modal(title,abstract)

AI总结 GrandTour数据集为多模态感知与状态估计提供了大规模野外腿式机器人数据,支持SLAM和高精度状态估计的研究与开发。

Comments Turcan Tuna, and Jonas Frey contributed equally. Submitted to Sage The International Journal of Robotics Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07708 2026-07-09 cs.CL cs.AI cs.CE cs.LG 新提交 62%

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

通过深度原生结构推理实现准确、跨学科和透明的结构-属性理解

Chen Tang, Yizhou Wang, Jianyu Wu, Lintao Wang, Shixiang Tang, Pengze Li, Encheng Su, Jun Yao, Jiabei Xiao, Yuqi Shi, Jielan Li, Hongxia Hao, Zhangyang Gao, Fang Wu, Ben Fei, Xiangyu Yue, Pan Tan, Bozitao Zhong, Jinouwen Zhang, Aoran Wang, Yan Lu, Jiaheng Liu, Xinzhu Ma, Liang Hong, Mingyue Zheng, Phil Torr, Bowen Zhou, Wanli Ouyang, Lei Bai

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) University of Sydney(悉尼大学) Nanjing University(南京大学) University of Oxford(牛津大学) The University of Science and Technology of China(中国科学技术大学) Drug Discovery and Design Center, State Key Laboratory of Drug Research, Shanghai Institute of Materia Medica, Chinese Academy of Sciences(药物发现与设计中心、国家药物研究重点实验室、上海中医药材料医学研究所、中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Stanford University(斯坦福大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 研究聚焦利用人工智能解释结构-属性关系的挑战,提出多模态科学基础模型SciReasoner,通过离散化结构信息为可寻址单元进行推理,在多领域基准测试中表现出色,实现准确预测与可解释科学推理的结合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17298 2026-07-09 cs.CV cs.AI 版本更新 62%

Explain Before You Answer: A Survey on Compositional Visual Reasoning

回答之前先解释:组合视觉推理综述

Fucai Ke, Joy Hsu, Zhixi Cai, Zixian Ma, Xin Zheng, Xindi Wu, Sukai Huang, Weiqing Wang, Pari Delir Haghighi, Gholamreza Haffari, Ranjay Krishna, Jiajun Wu, Hamid Rezatofighi

机构 * Monash University(墨尔本大学) Stanford University(斯坦福大学) University of Washington(华盛顿大学) Griffith University(格里菲斯大学) Princeton University(普林斯顿大学) Allen Institute for Artificial Intelligence(人工智能研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 综述2023年至2025年组合视觉推理文献,形式化核心定义,追溯范式转变,编目基准指标,提炼见解、识别挑战并概述方向,为该领域提供统一分类、历史路线图和批判性展望。

Comments Project Page: https://github.com/pokerme7777/Compositional-Visual-Reasoning-Survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07189 2026-07-09 cs.AI 新提交 57%

Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

人工智能理解成像吗?计算成像任务中智能体人工智能的系统基准测试

Ethan Chung, Chuanjun Zheng, Jasper Tan, Jingxi Li, Haopeng Zhang, Huaijin Chen

机构 * University of Hawaii at Manoa(夏威夷大学马诺阿分校) Glass Imaging(玻璃成像公司)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 研究人工智能在计算成像任务中的表现,提出ImagingBench基准测试,涵盖五类20个任务及三种评估设置,对多模态系统测试发现智能体模型比专门方法弱,该基准可测差距与跟踪进展。

Comments 14 pages, 11 figures. Preprint / work in progress. Paper Webpage: https://cirp-lab.github.io/imagingbench

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07179 2026-07-09 cs.CV cs.LG 新提交 57%

Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering

用于通用文档视觉问答的领域适应视觉语言模型的比较研究

Miguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno, Ruben Vera-Rodriguez, Daniel DeAlcala, Aythami Morales, Ruben Tolosana, Oscar Delgado, Alvaro Ortigosa, Javier Ortega-Garcia

机构 * Universidad Autónoma de Madrid (UAM)(马德里自治大学) BiometricsAI(生物识别人工智能)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 研究对8个开源预训练VLMs在三种文档领域的DocVQA进行全面评估,通过多种评估方式发现其在不同布局性能有差异,参数缩放影响性能,视觉理解是瓶颈,还表明少样本微调能让模型快速适应目标域文档。

Comments 17 pages, 4 figures, accepted at the Automatically Domain-Adapted and Personalized Document Analysis workshop of the ICDAR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06929 2026-07-09 cs.SD cs.AI 新提交 57%

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

MADB:一个具有专业和多维度注释的大规模音乐美学数据集

Sirui Zhang, Tianle Wang, Xinyi Tong, Peiyang Yu, Jishang Chen, Liangke Zhao, Haoxin Zhang, Duo Xu, Xin Jin, Feng Yu, Songchun Zhu

机构 * Central Conservatory of Music, China(中央音乐学院) Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院) Tianjin Conservatory of Music(天津音乐学院) Beijing Electronic Science and Technology Institute(北京电子科技研究所) Peking University(北京大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 研究音乐美学评估问题,引入含9999首曲目、由30名注释者标注的MADB数据集,通过多预训练模型建立统一评估框架,揭示模型与人类判断差距,为音乐理解提供新基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06838 2026-07-09 cs.CV 新提交 57%

WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence

WildCity:用于渲染、模拟和空间智能的真实世界城市规模测试平台

Xiangyu Han, Mengyu Yang, Jiaqi Li, Bowen Chang, Ziyu Chen, Hexu Zhao, Rahul Kumar Agrawal, Anthony Rodriguez, Fiona Hua, Marco Pavone, Chen Feng, Yiming Li

机构 * May Mobility(五月出行公司) New York University(纽约大学) NVIDIA(英伟达公司) Stanford University(斯坦福大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 针对人工智能在构建城市规模空间表征方面的挑战,引入WildCity真实世界多模态数据集,建立重建基线并转化为模拟器,分析关键挑战,推动城市规模渲染及相关人工智能发展。

Comments ECCV 2026; Project Page: https://han-xiangyu.github.io/Wild-City/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06691 2026-07-09 cs.CV 新提交 57%

CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views

CoMind:从多视角理解人类协作活动

Alexey Gavryushin, Dingxi Zhang, Zhao Huang, Alexandros Delitzas, Jiaqi Chen, Ben Ellis, Cedric Zöllner, Manthan Patel, Manuel Kaufmann, Marc Pollefeys, Xi Wang

机构 * ETH Zurich(苏黎世联邦理工学院) MPI for Informatics(马克斯·普朗克信息研究所) Microsoft Switzerland(微软瑞士公司) TU Munich(慕尼黑工业大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 研究旨在填补人类协作认知过程研究空白,引入CoMind数据集,整合多模态数据并标注,建立相关基准,有助于开发和评估能建模复杂社会互动及推理人类行为的AI系统。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07519 2026-07-09 cs.LG math.ST stat.TH 新提交 50%

Gradient-free Riemannian Langevin Sampler

无梯度黎曼朗之万采样器

Ricardo Baptista, Olivier Zahm

机构 * University of Toronto(多伦多大学) UGA, Inria, CNRS, Grenoble INP*, LJK(格勒诺布尔大学、法国国家信息与自动化研究所、法国国家科学研究中心、格勒诺布尔国立综合理工学院、数值模拟与知识工程实验室)

专题命中 多模态评测 :multimodal(abstract)

AI总结 研究多模态概率分布采样问题,提出无梯度黎曼朗之万采样器(GRiLS),通过引入黎曼度量改善探索,无需目标密度梯度评估,适用于复杂目标,用相互作用粒子系综估计均值和协方差,实证显示其混合效果优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06318 2026-07-09 cs.SI 版本更新 50%

The Schwurbelarchiv: a German Language Telegram dataset for the Study of Conspiracy Theories

Schwurbelarchiv:一个用于研究阴谋论的德语Telegram数据集

Mathias Angermaier, Elisabeth Hoeldrich, Jana Lasser, Joao Pinheiro Neto

专题命中 多模态评测 :multimodal(abstract)

AI总结 该研究构建了一个包含5800个群组和6300万条消息的Telegram数据集,通过解析和清洗原始数据,提供了研究德语阴谋论 discourse 的研究级资源。

Comments This paper is 20 pages, 2 figures, 4 tables, and one dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14590 2026-07-09 cs.LG 版本更新 50%

Counterfactual Modeling with Fine-Tuned LLMs for Health Intervention Design and Sensor Data Augmentation

基于微调大语言模型的反事实建模用于健康干预设计与传感器数据增强

Shovito Barua Soumma, Asiful Arefeen, Stephanie M. Carpenter, Melanie Hingle, Hassan Ghasemzadeh

机构 * College of Health Solutions, Arizona State University(亚利桑那州立大学健康解决方案学院) School of Computing and Augmented Intelligence, Arizona State University(亚利桑那州立大学计算与增强智能学院) School of Nutritional Sciences and Wellness, University of Arizona(亚利桑那大学营养科学与健康学院)

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文研究利用微调后的大型语言模型生成高质量反事实解释,用于健康干预设计和传感器数据增强,实验表明LLM生成的反事实能有效提升模型性能,具有临床应用价值。

Comments IEEE Open Journal of Engineering in Medicine and Biology (Volume: 7), Date of Publication: 28 May 2026. Page(s): 232-240

Journal ref IEEE Open Journal of Engineering in Medicine and Biology (Volume: 7), Date of Publication: 28 May 2026. Page(s): 232-240; Electronic ISSN: 2644-1276

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08366 2026-07-09 econ.GN q-fin.EC 版本更新 50%

A data fusion approach for mobility hub impact assessment and location selection: integrating hub usage data into a large-scale mode choice model

一种用于移动枢纽影响评估和选址的数据融合方法:将枢纽使用数据整合到大规模出行选择模型中

Xiyuan Ren, Joseph Y. J. Chow

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出一种数据融合方法,将移动枢纽使用数据整合到大规模出行选择模型中,评估枢纽对出行需求、出行方式转换、减少车辆行驶里程和增加消费者剩余的影响。

Journal ref Transportation Research Part A, 211 (2026), 105138

详情

展开后加载摘要…

URL PDF HTML 收藏