arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9080 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9080 篇

2509.23927 2026-01-26 cs.CV 88%

FUSAR-KLIP: Towards Multimodal Foundation Models for Remote Sensing

FUSAR-KLIP:迈向遥感多模态基础模型

Yi Yang, Xiaokun Zhang, Qingchen Fang, Jing Liu, Ziqi Ye, Rui Li, Li Liu, Haipeng Wang

机构 * Key Laboratory for Information Science of Electromagnetic Waves (MoE), Fudan University(电磁波信息科学重点实验室(MoE),复旦大学) Institute of Zhejiang Laboratory(浙江实验室研究院) College of Electronic Science and Technology, NUDT(电子科学与技术学院,南大学)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(title);cross-modal(abstract);分类 cs.CV

AI总结 FUSAR-KLIP是首个针对SAR图像的多模态基础模型,通过构建大规模数据集、生成结构化文本、设计自洽优化机制和建立统一评估基准,解决遥感图像与通用视觉表示之间的认知不一致问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13327 2025-08-20 cs.AI 88%

Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal Attention

Sarthak Khanna, Armin Berger, David Berghaus, Tobias Deusser, Lorenz Sparrenberg, Rafet Sifa

机构 * Fraunhofer IAIS - Department of Media Engineering(弗劳恩霍夫研究所媒体工程部门) University of Bonn - Department of Computer Science(波恩大学计算机科学系) West-AI - Federal Ministry of Education and Research(西德人工智能 - 教育与研究部)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

Comments Accepted in IEEE-DSAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21277 2025-05-22 cs.AI 88%

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

Guanghao Zhou, Panjia Qiu, Cen Chen, Jie Wang, Zheming Yang, Jian Xu, Minghui Qiu

机构 * East China Normal University(东华师范大学) ByteDance(字节跳动)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(title);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10875 2025-05-19 cs.CV 88%

A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision

Alexey Magay, Dhurba Tripathi, Yu Hao, Yi Fang

机构 * Embodied AI and Robotics (AIR) Lab New York University Abu Dhabi(embodied AI 和机器人(AIR)实验室 新 York 大学阿布扎赫德)

专题命中 多模态评测 :multimodal(title);multimodal foundation model(title);multi-modal(abstract);MLLM(abstract)

Comments Project website and code: https://dktpt44.github.io/LV-GPT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10462 2025-04-03 cs.CV 88%

CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Wei Chen, Lin Li, Yongqi Yang, Bin Wen, Fan Yang, Tingting Gao, Yu Wu, Long Chen

专题命中 多模态评测 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CV

Comments 22 pages, Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01035 2025-03-17 cs.CV cs.LG 88%

Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality

Hu Wang, Congbo Ma, Jianpeng Zhang, Yuan Zhang, Jodie Avery, Louise Hull, Gustavo Carneiro

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV

Journal ref Medical Image Computing and Computer-Assisted Intervention 2023 (MICCAI 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06885 2025-03-11 cs.CV 88%

ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks

Yan Yang, Dongxu Li, Haoning Wu, Bei Chen, Liu Liu, Liyuan Pan, Junnan Li

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(title);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02621 2024-12-04 cs.AI cs.LG 88%

Medical Multimodal Foundation Models in Clinical Diagnosis and Treatment: Applications, Challenges, and Future Directions

Kai Sun, Siyan Xue, Fuchun Sun, Haoran Sun, Yu Luo, Ling Wang, Siyuan Wang, Na Guo, Lei Liu, Tian Zhao, Xinzhou Wang, Lei Yang, Shuo Jin, Jun Yan, Jiahong Dong

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13264 2024-10-14 cs.AI cs.LG cs.SE 88%

WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks

Michael Wornow, Avanika Narayan, Ben Viggiano, Ishan S. Khare, Tathagat Verma, Tibor Thompson, Miguel Angel Fuentes Hernandez, Sudharsan Sundar, Chloe Trujillo, Krrish Chawla, Rongfei Lu, Justin Shen, Divya Nagaraj, Joshua Martinez, Vardhan Agrawal, Althea Hudson, Nigam H. Shah, Christopher Re

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13951 2024-09-17 cs.CL 88%

MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria

Wentao Ge, Shunian Chen, Guiming Hardy Chen, Junying Chen, Zhihong Chen, Nuo Chen, Wenya Xie, Shuo Yan, Chenghao Zhu, Ziyue Lin, Song Dingjie, Xidong Wang, Anningzhe Gao, Zhang Zhiyi, Jianquan Li, Xiang Wan, Benyou Wang

专题命中 多模态评测 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CL

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15802 2025-08-25 cs.CL cs.AI 88%

MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding

Mohan Jiang, Jin Gao, Jiahao Zhan, Dequan Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17489 2025-03-25 cs.CL cs.CV 88%

Judge Anything: MLLM as a Judge Across Any Modality

Shu Pu, Yaochen Wang, Dongping Chen, Yuhang Chen, Guohao Wang, Qi Qin, Zhongyi Zhang, Zhiyuan Zhang, Zetong Zhou, Shuang Gong, Yi Gui, Yao Wan, Philip S. Yu

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract);any-to-any(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21455 2026-07-24 cs.NI 新提交 88%

Out-of-Distribution Detection in Wireless Multimodal Foundation Models for 6G ISAC

6G智能感知与通信无线多模态基础模型中的分布外检测

Mohammad Farzanullah, Akram Bin Sediq, Ali Afana, Melike Erol-Kantarci

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(title,abstract)

AI总结 研究6G智能感知与通信中无线多模态基础模型的分布外检测问题,提出WMFM - OOD框架,通过构建基站原型和温度缩放概率评分机制区分异常,在DeepVerse6G数据集验证,显著优于基线,提升检测灵敏度,保障网络可靠性。

Comments presented at IEEE VTC 2026 Fall, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26312 2026-06-30 stat.ME 88%

Cross-modal dependence analysis with asynchronous longitudinal multimodal data

异步纵向多模态数据的跨模态依赖分析

Kun Qian, Hyung G. Park

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(title,abstract)

AI总结 提出一种贝叶斯潜变量模型,用于估计异步观测的多模态数据中协变量辅助的依赖结构,并应用于阿尔茨海默病神经影像学倡议数据,揭示临床有意义的纵向跨模态生物标志物依赖模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23402 2026-02-03 cs.LG stat.ML 88%

Quantized-Tinyllava: a new multimodal foundation model enables efficient split learning

量化-小TinyLLaVA:一种新的多模态基础模型实现了高效的分裂学习

Jiajun Guo, Xin Luo, Jiayin Zheng, Yiqun Wang, Kai-Wei Chang, Wei Wang, Jie Liu

机构 * Department of Statistics University of Michigan(统计学系密歇根大学) Department of Computational Medicine & Bioinformatics University of Michigan(计算医学与生物信息学系密歇根大学) Department of Biostatistics University of Michigan(生物统计学系密歇根大学) Department of Computer Science University of California, Los Angeles(计算机科学系加州大学洛杉矶分校)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(title,abstract)

AI总结 Quantized-TinyLLaVA通过量化压缩和高效分裂学习框架,在减少通信开销的同时保持模型性能,提升隐私保护能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23897 2026-01-01 cs.NI 88%

Wireless Multimodal Foundation Model (WMFM): Integrating Vision and Communication Modalities for 6G ISAC Systems

无线多模态基础模型(WMFM):为6G ISAC系统集成视觉与通信模态

Mohammad Farzanullah, Han Zhang, Akram Bin Sediq, Ali Afana, Melike Erol-Kantarci

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(title,abstract)

AI总结 本文提出基于对比学习的无线多模态基础模型WMFM,通过联合学习无线信道系数和视觉图像,实现6G ISAC系统的高效多模态学习与应用。

Comments Journal Paper, 13 pages, 11 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13092 2025-12-15 cs.LG cs.HC 88%

Uncertainty-Aware Cross-Modal Knowledge Distillation with Prototype Learning for Multimodal Brain-Computer Interfaces

具有原型学习的不确定性感知跨模态知识蒸馏用于多模态脑机接口

Hyo-Jeong Jang, Hye-Bin Shin, Seong-Whan Lee

机构 * Department of Brain and Cognitive Engineering, Korea University(脑科学与认知工程系,韩国大学) Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(title,abstract)

AI总结 本文提出一种具有原型学习的跨模态知识蒸馏框架,通过缓解模态和标签不一致问题,提升多模态脑机接口中EEG的分类和回归性能。

Comments Accepted to SMC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03464 2025-12-04 cs.LG 88%

Multi-Modal Opinion Integration for Financial Sentiment Analysis using Cross-Modal Attention

多模态意见整合用于金融情绪分析的跨模态注意力

Yujing Liu, Chen Yang

机构 * College of Computing Georgia Institute of Technology Atlanta, USA(计算学院 佐治亚理工学院 美国亚特兰大) College of Engineering University of Pennsylvania Philadelphia, USA(工程学院 宾夕法尼亚大学 美国费城)

专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(title);multimodal(abstract)

AI总结 本文提出了一种多模态注意力机制,用于整合金融意见的时效和流行度模态,以提高金融情绪分析的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01293 2025-06-03 cs.CV cs.AI cs.CL 88%

Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation

Yichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu, Min Zhang, Wen Zhang, Huajun Chen

机构 * Zhejiang University(浙江大学) Tianjin University(天津大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态评测 :multi-modal(title,abstract);MLLM(title);分类 cs.CV、cs.CL、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18170 2025-02-04 cs.LG 88%

Continually Evolved Multimodal Foundation Models for Cancer Prognosis

Jie Peng, Shuang Zhou, Longwei Yang, Yiran Song, Mohan Zhang, Kaixiong Zhou, Feng Xie, Mingquan Lin, Rui Zhang, Tianlong Chen

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(title);multi-modal(abstract)

Comments 9 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20775 2024-08-22 cs.CR cs.AI cs.CL cs.MM 88%

Medical MLLM is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models

Xijie Huang, Xinyuan Wang, Hantao Zhang, Yinghao Zhu, Jiawen Xi, Jingkun An, Hao Wang, Hao Liang, Chengwei Pan

专题命中 多模态评测 :multimodal(title,abstract);MLLM(title);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15486 2023-03-29 cs.LG 88%

Unimodal Training-Multimodal Prediction: Cross-modal Federated Learning with Hierarchical Aggregation

Rongyu Zhang, Xiaowei Chi, Guiliang Liu, Wenyi Zhang, Yuan Du, Fangxin Wang

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(title,abstract)

Comments 10 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02241 2021-08-06 cs.LG eess.SP 88%

Attentive Cross-modal Connections for Deep Multimodal Wearable-based Emotion Recognition

Anubhav Bhatti, Behnam Behinaein, Dirk Rodenburg, Paul Hungler, Ali Etemad

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(title,abstract)

Comments 5 pages, 2 figures. Accepted at 2021 9th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09281 2026-08-11 cs.AI 新提交 87%

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

MMArch:基于建筑证据的多模态推理基准测试

Chenxu Du, Kang An, Tengyue Wang, Zhongyu Yang, Xinqi Yang, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren

专题命中 多模态评测 :multimodal(title,abstract);MLLM(summary_cn,abstract_cn);分类 cs.AI

AI总结 研究针对多模态大语言模型(MLLM)在工程多模态推理基准上的表现,构建MMArch基准,发现现有MLLM与人类专家存在差距,为相关研究提供评估基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03078 2026-08-05 cs.CV 新提交 87%

LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds

LDU-Bench:面向不同布局电路背景下光刻缺陷理解的多模态大语言模型评估基准

Huanglong Ji, Botong Zhao, Shujing Lv, Yue Lv

专题命中 多模态评测 :multimodal(title,abstract);MLLM(summary_cn);分类 cs.CV

AI总结 本文提出LDU-Bench这一多模态基准,将光刻审查工作流程分解为四项任务,评估发现现有多模态大语言模型的缺陷分类能力无法稳定迁移至下游审查阶段,为工业MLLM评估提供了统一平台。

Comments 12 pages, 3 figures, and 5 tables, including appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01849 2026-08-04 cs.AI 新提交 87%

Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models

探索与弥合未学习多模态大语言模型中的知识缺口

Junxiang You, Junkai Chen, Yuhao He, Ruiqi Liu, Zhetao Guo, Shu Wu

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.AI

AI总结 针对未学习多模态大语言模型存在的知识缺口问题,提出带锚定正则化的选择性保护方法,在保障未学习安全性的同时大幅恢复模型响应质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27637 2026-08-04 cs.CV 版本更新 87%

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

MMOOC:多模态大语言模型的上下文外评估综合基准

Wenjie Zhu, Yabin Zhang, Wenjun Zeng, Lei Zhang

机构 * The Hong Kong Polytechnic University(香港理工大学) Eastern Institute of Technology(东方理工大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出MMOOC基准,评估多模态大语言模型在上下文偏移场景下的拒绝与作答能力,实验发现当前模型难以平衡可答性与拒绝性,后训练可提升稳健性,该基准将公开。

Comments project page:https://zhuwenjie98.github.io/MMOOC-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27667 2026-07-31 cs.CV 新提交 87%

Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers

证据见证组合:针对闭集多模态答案的单预填充风险检测

Fexiang Liu, Shiye Wang, Qiang Qiu, Zheng Wang

专题命中 多模态评测 :multimodal(title,abstract);MLLM(summary_cn,abstract_cn);分类 cs.CV

AI总结 该研究提出WEP方法,利用MLLM的白盒预填充路径,无需额外操作即可检测闭集视觉答案的推理风险,在3个MLLM和4个基准上提升了平均错误AP。

Comments 22 pages, 6 figures; includes supplementary material. Code: https://github.com/SouthWinter/WEP

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24957 2026-07-29 cs.CV 新提交 87%

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

感知基准:评估多模态大语言模型中的原子视觉感知

Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng Gao, Lin Sui, Weihong Li, Xinxing Zu, Jia Chen, Yao Wang, Xiaoxue Wu, Yalin Wang, Y. Charles, Yiping Bao, Yangyang Liu, Zhiqi Huang, Xinyu Zhou

机构 * Moonshot AI(登月人工智能)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(summary_cn,abstract_cn);分类 cs.CV

AI总结 介绍用于评估多模态大语言模型原子视觉感知能力的PerceptionBench基准,通过自下而上方法构建错误分类法及相关问题,测试16个前沿模型,发现原子感知待解决,该基准为衡量MLLM视觉感知边界提供标准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08839 2026-07-13 cs.CV cs.LG 新提交 87%

Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing

混合探针:通过探针在多模态大语言模型中从特权模态学习

Dominick Reilly, Qiyu Wu, Hiromi Wakaki, Srijan Das, Yuki Mistufuji

机构 * Sony Group Corporation(索尼集团公司) Sony AI(索尼人工智能) University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract,abstract_cn);cross-modal(abstract);分类 cs.CV

AI总结 研究多模态大语言模型在特权模态设置下的问题,提出混合探针(MoP)框架及MoP跨模态训练(MoP-X),通过结构化探测机制分离模态信号,经评估在多任务中优于基线,有效利用辅助模态训练可提升性能。

Comments Preprint (16 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏