arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9087 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9087 篇

2509.23879 2025-09-30 cs.CV cs.AI cs.CL cs.MM 85%

PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications

Hitesh Laxmichand Patel, Amit Agarwal, Srikant Panda, Hansa Meghwani, Karan Dua, Paul Li, Tao Sheng, Sujith Ravi, Dan Roth

机构 * Oracle AI

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted in EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03823 2025-09-23 cs.CV cs.AI cs.CL cs.MM 85%

Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM

Dingjie Song, Sicheng Lai, Mingxuan Wang, Shunian Chen, Lichao Sun, Benyou Wang

机构 * Lehigh University(莱维大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08800 2025-09-11 cs.SD cs.AI cs.CV cs.MM eess.AS 85%

PianoVAM: A Multimodal Piano Performance Dataset

Yonghyun Kim, Junhyung Park, Joonhyung Bae, Kirak Kim, Taegyun Kwon, Alexander Lerch, Juhan Nam

专题命中 多模态评测 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to the 26th International Society for Music Information Retrieval (ISMIR) Conference, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02429 2025-08-05 cs.AI cs.LG 85%

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

Miaosen Luo, Jiesen Long, Zequn Li, Yunying Yang, Yuncheng Jiang, Sijie Mai

机构 * School of Computer Science, South China Normal University(华南师范大学计算机科学学院) School of Information Technology in Education, South China Normal University(华南师范大学教育信息技术学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16206 2025-07-23 cs.LG cs.AI 85%

METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark

Xu Yang, Qi Zhang, Shuming Jiang, Yaowen Xu, Zhaofan Zou, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(电信人工智能研究院) Institute of Artificial Intelligence and Robotics(IAIR), Xi’an Jiaotong University(人工智能与机器人研究院) Advanced Technique of Artificial Intelligence(ATAI), Chongqing University of Technology(人工智能先进技术研究院)

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.AI

Comments 9 pages,3 figures ICCV format

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14449 2025-07-22 cs.CV 85%

IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark

Zhe Cao, Jin Zhang, Ruiheng Zhang

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments 11 pages, 7 figures. This paper is accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20104 2025-06-16 cs.CV 85%

New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration

Xuzheng Yang, Junzhuo Liu, Peng Wang, Guoqing Wang, Yang Yang, Heng Tao Shen

机构 * University of Electronic Science and Technology of China(电子科技大学) Tongji University(同济大学)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09745 2025-06-12 cs.CV 85%

Class Similarity-Based Multimodal Classification under Heterogeneous Category Sets

Yangrui Zhu, Junhua Bao, Yipan Wei, Yapeng Li, Bo Du

机构 * School of Computer Science Wuhan University(武汉大学计算机学院)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14705 2025-05-22 cs.CV cs.LG 85%

Beyond Modality Collapse: Representations Blending for Multimodal Dataset Distillation

Xin Zhang, Ziruo Zhang, Jiawei Du, Zuozhu Liu, Joey Tianyi Zhou

机构 * Centre for Frontier AI Research, Agency for Science, Technology and Research, Singapore(前沿人工智能研究中心,科技研究局,新加坡) Institute of High Performance Computing, Agency for Science, Technology and Research, Singapore(高性能计算研究所,科技研究局,新加坡) National University of Singapore, Singapore(新加坡国立大学) Zhejiang University, China(浙江大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08813 2025-04-15 cs.LG cs.AI cs.CR 85%

SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models

Junfeng Fang, Yukai Wang, Ruipeng Wang, Zijun Yao, Kun Wang, An Zhang, Xiang Wang, Tat-Seng Chua

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14941 2025-03-20 cs.CV 85%

UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation

Qihui Zhang, Munan Ning, Zheyuan Liu, Yanbo Wang, Jiayi Ye, Yue Huang, Shuo Yang, Xiao Chen, Yibing Song, Li Yuan

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06777 2024-10-10 cs.CV 85%

HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding

Keliang Li, Zaifei Yang, Jiahe Zhao, Hongze Shen, Ruibing Hou, Hong Chang, Shiguang Shan, Xilin Chen

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02608 2024-09-05 cs.CV 85%

A Medical Multimodal Large Language Model for Pediatric Pneumonia

Weiwei Tian, Xinyu Huang, Tianhao Cheng, Wen He, Jinwu Fang, Rui Feng, Daoying Geng, Xiaobo Zhang

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

Comments 18 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03701 2024-06-12 cs.MM 85%

Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction

Meishan Zhang, Hao Fei, Bin Wang, Shengqiong Wu, Yixin Cao, Fei Li, Min Zhang

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17092 2023-11-30 cs.CV 85%

SEED-Bench-2: Benchmarking Multimodal Large Language Models

Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, Ying Shan

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

Comments Project released at: https://github.com/AILab-CVC/SEED-Bench. arXiv admin note: text overlap with arXiv:2307.16125

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05523 2023-11-01 cs.CL cs.AI cs.CV cs.MM 85%

FACTIFY3M: A Benchmark for Multimodal Fact Verification with Explainability through 5W Question-Answering

Megha Chakraborty, Khushbu Pahwa, Anku Rani, Shreyas Chatterjee, Dwip Dalal, Harshit Dave, Ritvik G, Preethi Gurumurthy, Adarsh Mahor, Samahriti Mukherjee, Aditya Pakala, Ishan Paul, Janvita Reddy, Arghya Sarkar, Kinjal Sensharma, Aman Chadha, Amit P. Sheth, Amitava Das

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments arXiv admin note: text overlap with arXiv:2305.04329

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12626 2023-09-26 cs.AI 85%

Enhancing Human-like Multi-Modal Reasoning: A New Challenging Dataset and Comprehensive Framework

Jingxuan Wei, Cheng Tan, Zhangyang Gao, Linzhuang Sun, Siyuan Li, Bihui Yu, Ruifeng Guo, Stan Z. Li

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.06767 2022-09-30 cs.CV cs.LG 85%

Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark

Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou, Minzhe Niu, Xiaodan Liang, Lewei Yao, Runhui Huang, Wei Zhang, Xin Jiang, Chunjing Xu, Hang Xu

专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2022 Track Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.03403 2020-10-08 cs.CV 85%

Universal Weighting Metric Learning for Cross-Modal Matching

Jiwei Wei, Xing Xu, Yang Yang, Yanli Ji, Zheng Wang, Heng Tao Shen

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15296 2024-12-10 cs.CV cs.AI cs.CL 85%

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Chaoyou Fu, Yi-Fan Zhang, Shukang Yin, Bo Li, Xinyu Fang, Sirui Zhao, Haodong Duan, Xing Sun, Ziwei Liu, Liang Wang, Caifeng Shan, Ran He

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Produced by MME+MMBench+LLaVA Teams. Project Page: https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13549 2024-12-02 cs.CV cs.AI cs.CL cs.LG 85%

A Survey on Multimodal Large Language Models

Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, Enhong Chen

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted for publication in National Science Review. Project page:https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13566 2023-09-12 cs.LG cs.AI cs.CL cs.CV 85%

MLLM-DataEngine: An Iterative Refinement Approach for MLLM

Zhiyuan Zhao, Linke Ouyang, Bin Wang, Siyuan Huang, Pan Zhang, Xiaoyi Dong, Jiaqi Wang, Conghui He

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Code and models are available at https://github.com/opendatalab/MLLM-DataEngine

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.09253 2022-03-29 cs.CV cs.CL cs.MM 85%

Logically at Factify 2022: Multimodal Fact Verification

Jie Gao, Hella-Franziska Hoffmann, Stylianos Oikonomou, David Kiskovski, Anil Bandhakavi

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted in AAAI'22: First Workshop on Multimodal Fact-Checking and Hate Speech Detection, Februrary 22 - March 1, 2022,Vancouver, BC, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18121 2026-08-11 cs.CV cs.AI 版本更新 85%

VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging

VCU-Bridge: 通过语义桥梁实现层次化视觉隐喻理解

Ming Zhong, Yuanlei Wang, Liuzhou Zhang, Ruichuan An, Renrui Zhang, Hao Liang, Ming Lu, Ying Shen, Wentao Zhang

机构 * Zhejiang University(浙江大学) Peking University(北京大学) Sun Yat-sen University(中山大学) CUHK(香港中文大学)

专题命中 多模态评测 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 VCU-Bridge通过语义桥梁实现层次化视觉隐喻理解,构建了HVCU-Bench基准,并展示了提升MLLM能力的显著效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02059 2026-08-04 cs.CV cs.MM 新提交 85%

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

MIEScore:面向多源图像编辑的人类对齐评估方法

Zitong Xu, Huiyu Duan, Xinyun Zhang, Weifei Xiong, Tianyi Zheng, Xiongkuo Min, Qiang Hu, Zhengxue Cheng, Bo Li, Guangtao Zhai

专题命中 多模态评测 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV、cs.MM

AI总结 针对多源图像编辑(MIE)缺乏人类对齐评估基准的问题,研究构建了首个MIE基准MIE-Bench,并提出基于MLLM的评估模型MIEScore,其对齐人类偏好性能最优且泛化性良好。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15749 2026-06-16 cs.CV cs.AI cs.SY eess.SY 新提交 85%

OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

OmniTraffic:面向时空交通推理的可控生成流水线与基准

Maonan Wang, Zhengyan Huang, Kemou Jiang, Yuhang Fu, Jiayue Zhu, Yuxin Cai, Xingchen Zou, Qiaosheng Zhang, Yi Yu, Ding Wang, Xi Chen, Ben M. Chen, Yuxuan Liang, Zhiyong Cui, Man On Pun, Yirong Chen

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shanghai AI Lab(上海人工智能实验室) Beihang University(北京航空航天大学) Nanyang Technological University(南洋理工大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态评测 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出OmniTraffic,一个基于12个真实路口3D重建的可控生成流水线与基准,通过8M VQA样本和3K人工验证测试集评估11个前沿MLLM,揭示拓扑与时空推理中的显著人机差距,并证明仿真数据微调可提升真实场景性能。

Comments 34 pages, 28 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07613 2026-06-09 cs.CV cs.AI 新提交 85%

Can You Trust What You See? Human and AI Detection of Synthetic Legal Evidence

你能相信你所见的吗?人类与AI对合成法律证据的检测

Jinzhe Tan, Ali Ekber Cinar, Karim Benyekhlef

机构 * Faculty of Law, McGill University(麦吉尔大学法学院)

专题命中 多模态评测 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 研究人类和前沿多模态大模型在民事纠纷场景中区分真实照片与AI生成图像的能力,发现两者均不可靠,提出结合人工审查、MLLM筛查和来源认证的解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00148 2026-06-02 cs.CV cs.AI 85%

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

StemBind: 当多模态大语言模型在抽象视觉推理中迷失于规则与实例之间

Xixiang He, Baiqi Wu, Xingming Li, Ao Cheng, Qiyao Sun, Xuanyu Ji, Qingyong Hu

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态评测 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出 StemBind 诊断基准,通过共享主干的三对齐问题(感知、规则、完整)定位 MLLM 在抽象视觉推理中的失败环节,发现规则到实例的绑定是主要瓶颈。

Comments Project page: https://hexixiang.github.io/StemBind

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14830 2025-10-10 cs.CV cs.AI cs.LG 85%

ProtoMedX: Towards Explainable Multi-Modal Prototype Learning for Bone Health Classification

Alvaro Lopez Pellicer, Andre Mariucci, Plamen Angelov, Marwan Bukhari, Jemma G. Kerns

机构 * School of Computing and Communications(计算与通讯学院) Lancaster Medical School(兰卡斯特医学学院)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract,comments);分类 cs.CV、cs.AI

Comments ICCV 2025 (PHAROS-AFE-AIMI: Adaptation, Fairness, and Explainability in Medical Imaging). 8 pages, 5 figures, 4 tables. Keywords: multi-modal, multimodal, prototype learning, explainable AI, interpretable models, case-based reasoning, medical imaging, DEXA, bone health, osteoporosis, osteopenia, diagnosis, classification, clustering

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01266 2024-08-20 cs.AI cs.CL 85%

IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations

Deqing Fu, Ruohao Guo, Ghazal Khalighinejad, Ollie Liu, Bhuwan Dhingra, Dani Yogatama, Robin Jia, Willie Neiswanger

专题命中 多模态评测 :multimodal(title);multimodal foundation model(title);分类 cs.CL、cs.AI

Comments 1st Conference on Language Modeling (COLM), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏