arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9087 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9087 篇

2205.10237 2022-05-23 cs.CL cs.AI 84%

M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue Database

Jinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu, Qin Jin, Xinchao Wang, Haizhou Li

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CL、cs.AI

Journal ref published at ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.07316 2022-05-04 cs.CL cs.AI cs.LG 84%

XDBERT: Distilling Visual Information to BERT from Cross-Modal Systems to Improve Language Understanding

Chan-Jan Hsu, Hung-yi Lee, Yu Tsao

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL、cs.AI

Comments ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12692 2022-04-01 cs.MM cs.CV 84%

Affective Feedback Synthesis Towards Multimodal Text and Image Data

Puneet Kumar, Gaurav Bhat, Omkar Ingle, Daksh Goyal, Balasubramanian Raman

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.MM

Comments Submitted to ACM Transactions on Multimedia Computing, Communications, and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.05834 2022-01-19 cs.CV cs.MM 84%

Tailor Versatile Multi-modal Learning for Multi-label Emotion Recognition

Yi Zhang, Mingyuan Chen, Jundong Shen, Chongjun Wang

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments To be published in AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08470 2021-12-17 cs.CL cs.CV 84%

Insta-VAX: A Multimodal Benchmark for Anti-Vaccine and Misinformation Posts Detection on Social Media

Mingyang Zhou, Mahasweta Chakraborti, Sijia Qian, Zhou Yu, Jingwen Zhang

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.06053 2021-10-22 cs.MM cs.CL 84%

The Multimodal Sentiment Analysis in Car Reviews (MuSe-CaR) Dataset: Collection, Insights and Improvements

Lukas Stappen, Alice Baird, Lea Schumann, Björn Schuller

专题命中 多模态评测 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CL、cs.MM

Comments accepted version

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04529 2021-04-07 cs.CV cs.AI 84%

Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting

Lingbo Liu, Jiaqi Chen, Hefeng Wu, Guanbin Li, Chenglong Li, Liang Lin

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by CVPR2021. Our code and benchmark for RGBT crowd counting are released at {\url{http://lingboliu.com/RGBT_Crowd_Counting.html}}

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.01513 2021-01-06 cs.CV cs.AI 84%

Deep Class-Specific Affinity-Guided Convolutional Network for Multimodal Unpaired Image Segmentation

Jingkun Chen, Wenqi Li, Hongwei Li, Jianguo Zhang

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.03240 2019-08-06 cs.MM cs.CL 84%

Informative Visual Storytelling with Cross-modal Rules

Jiacheng Li, Haizhou Shi, Siliang Tang, Fei Wu, Yueting Zhuang

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL、cs.MM

Comments 9 pages, to appear in ACM Multimedia 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.06462 2018-08-27 cs.CY cs.AI cs.MM q-bio.QM 84%

Cross-Modal Health State Estimation

Nitish Nag, Vaibhav Pandey, Preston J. Putzel, Hari Bhimaraju, Srikanth Krishnan, Ramesh C. Jain

专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI、cs.MM

Comments Accepted to ACM Multimedia 2018 Conference - Brave New Ideas, Seoul, Korea, ACM ISBN 978-1-4503-5665-7/18/10

Journal ref Nitish Nag, Vaibhav Pandey, Preston J. Putzel, Hari Bhimaraju, Srikanth Krishnan, Ramesh C. Jain, 2018 ACM Multimedia Conference (MM '18), October 22--26, 2018, Seoul, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.05147 2018-06-15 cs.CV cs.MM 84%

Cross-modal Hallucination for Few-shot Fine-grained Recognition

Frederik Pahde, Patrick Jähnichen, Tassilo Klein, Moin Nabi

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

Comments CVPR 2018 Workshop on Fine-Grained Visual Categorization

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09644 2026-04-14 cs.CY cs.AI 84%

Detecting Corporate AI-Washing via Cross-Modal Semantic Inconsistency Learning

通过跨模态语义不一致性学习检测企业AI洗钱

Zhanjie Wen, Jingqiao Guo

机构 * School of Economics and Trade, Guangdong University of Finance(广东金融学院经济贸易学院) Department of Computer Science, Faculty of Science, Hong Kong Baptist University(香港浸会大学理学院计算机科学系)

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract,comments);分类 cs.AI

AI总结 本文提出AWASH框架,通过跨模态主张-证据推理检测企业AI洗钱,利用AW-Bench基准测试,实现高准确率的AI能力识别。

Comments 28 pages, 6 figures, Journal Submission (Finance/Accounting & Computer Science Interdiscipline), 6 tables, 40 references, trimodal benchmark (88,412 firm-quarter observations) and end-to-end multimodal detection framework for corporate AI-washing

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03437 2026-03-05 cs.CV 84%

Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning

超越准确率:评估多模态医学推理中的视觉语义

Anas Zafar, Leema Krishna Murali, Ashish Vashist

机构 * The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心) Cohere Labs(Cohere实验室) Eisai Inc.(艾伯维公司) Indian Institute of Science, Bangalore(班加罗尔印度科学研究院)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 该研究提出反事实评估框架,通过测量视觉依赖性得分和幻觉视觉推理率,揭示仅文本强化学习在多模态医学推理中降低视觉依赖性的问题。

Comments 12 pages, 2 figures, 2 tables, medical VQA / multimodal reasoning evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18991 2025-12-09 cs.CL 84%

Surveying the MLLM Landscape: A Meta-Review of Current Surveys

测绘MLLM景观:当前调研的元综述

Ming Li, Keyu Chen, Ziqian Bi, Ming Liu, Xinyuan Song, Zekun Jiang, Tianyang Wang, Benji Peng, Qian Niu, Junyu Liu, Jinlang Wang, Sen Zhang, Xuanhe Pan, Jiawei Xu, Pohsun Feng

机构 * Georgia Institute of Technology(佐治亚理工学院) Indiana University(印第安纳大学) Purdue University(普渡大学) Emory University(埃默里大学) Sichuan University(四川大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) AppCubic Kyoto University(京都大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Rutgers University(罗格斯大学) National Taiwan Normal University(台湾师范大学)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract,comments);分类 cs.CL

AI总结 本文通过系统回顾现有调研,全面分析MLLMs的评估方法、挑战与趋势,为研究者提供当前评估现状的深入理解。

Comments The article consists of 22 pages, including 2 figures and 108 references. The paper provides a meta-review of surveys on Multimodal Large Language Models (MLLMs), categorizing findings into key areas such as evaluation, applications, security, and future directions

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00082 2025-12-02 cs.CV 84%

Exploring Diagnostic Prompting Approach for Multimodal LLM-based Visual Complexity Assessment: A Case Study of Amazon Search Result Pages

探索用于多模态大语言模型视觉复杂性评估的诊断提示方法:亚马逊搜索结果页面案例研究

Divendar Murtadak, Yoon Kim, Trilokya Akula

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本研究通过对比诊断提示与传统提示方法,发现其在提升多模态大语言模型对亚马逊搜索结果页面视觉复杂性评估的可靠性方面有显著改进,但仍有待进一步优化。

Comments 9 pages, 4 figures, 9 tables. Study on diagnostic prompting for multimodal LLM-based visual complexity assessment of Amazon search result pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13394 2025-10-27 cs.CV 84%

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, Rongrong Ji, Caifeng Shan, Ran He

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) Tencent Youtu Lab(腾讯优图实验室) Xiamen University(厦门大学) CASIA(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments NeurIPS DB 2025 Spotlight, Project Page: https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21635 2025-09-01 cs.RO cs.CV cs.SY eess.SY 84%

The Rosario Dataset v2: Multimodal Dataset for Agricultural Robotics

Nicolas Soncini, Javier Cremona, Erica Vidal, Maximiliano García, Gastón Castro, Taihú Pire

机构 * CIFASIS (CONICET-UNR)(CIFASIS(CONICET-UNR)) Universidad de San Andrés (UDESA-CONICET)(Universidad de San Andrés(UDESA-CONICET))

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract,journal_ref);分类 cs.CV

Comments First published on The International Journal of Robotics Research: https://journals.sagepub.com/doi/10.1177/02783649251368909

Journal ref The Rosario dataset v2: Multi-modal dataset for agricultural robotics. The International Journal of Robotics Research. 2025;0(0)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22041 2025-06-30 eess.IV cs.CV 84%

Towards Scalable and Robust White Matter Lesion Localization via Multimodal Deep Learning

Julia Machnio, Sebastian Nørgaard Llambias, Mads Nielsen, Mostafa Mehdipour Ghazi

机构 * Pioneer Centre for AI(先锋人工智能中心) University of Copenhagen(哥本哈根大学)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract,comments);分类 cs.CV

Comments 2nd Sorbonne-Heidelberg Workshop on AI in medicine: Machine Learning for multi-modal data

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11080 2025-06-16 cs.CL 84%

MANBench: Is Your Multimodal Model Smarter than Human?

Han Zhou, Qitong Xu, Yiheng Dong, Xin Yang

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(电子信息与通讯学院,华中科技大学) School of Basic Medicine, Huazhong University of Science and Technology(基础医学院,华中科技大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Multimodal Benchmark, Project Url: https://github.com/micdz/MANBench, ACL2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04781 2025-04-08 cs.CV 84%

OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance

Chaoyi Wang, Baoqing Li, Xinhan Di

专题命中 多模态评测 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV;multimodal(comments)

Comments This work has been accepted to the Multimodal Algorithmic Reasoning (MAR) Workshop at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16860 2024-12-05 cs.CV 84%

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Ziteng Wang, Rob Fergus, Yann LeCun, Saining Xie

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract,comments);分类 cs.CV

Comments NeurIPS 2024 (Oral). Website at https://cambrian-mllm.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07053 2024-10-04 cs.CV 84%

Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model

Wenqi Zhang, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu, Yueting Zhuang

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract,comments);分类 cs.CV

Comments The paper is accepted by EMNLP-24. Code: https://github.com/zwq2018/Multi-modal-Self-instruct dataset: https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct Leaderboard: https://multi-modal-self-instruct.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00578 2024-04-02 cs.CV 84%

M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Fan Bai, Yuxin Du, Tiejun Huang, Max Q. -H. Meng, Bo Zhao

专题命中 多模态评测 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV;MLLM(comments)

Comments MLLM, 3D medical image analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02556 2023-04-06 cs.CV 84%

Detecting and Grounding Multi-Modal Media Manipulation

Rui Shao, Tianxing Wu, Ziwei Liu

专题命中 多模态评测 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV;multimodal(comments)

Comments CVPR 2023. Project page: https://rshaojimmy.github.io/Projects/MultiModal-DeepFake Code: https://github.com/rshaojimmy/MultiModal-DeepFake

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01371 2026-05-05 cs.RO 83%

ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue

ESARBench: 一种用于代理无人机体感知搜索与救援的基准

Daoxuan Zhang, Ping Chen, Jianyi Zhou, Shuo Yang

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态评测 :MLLM(summary_cn,abstract);multimodal(abstract)

AI总结 本文提出ESARBench,首个用于评估基于MLLM的无人机在高真实感搜索与救援场景中的基准,通过构建高保真环境和600个任务数据集,揭示了ESAR中的关键瓶颈。

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09648 2026-08-13 cs.DB cs.AI 版本更新 83%

ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset

ArtiFact: 大规模多模态文化遗产数据集

Luciano Duarte, Olga Ovcharenko, Sebastian Schelter

机构 * BIFOLD & TU Berlin(BIFOLD与柏林技术大学)

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 提出包含65万条博物馆记录的多模态文化遗产数据集ArtiFact,用于跨模态错误检测和语义查询处理,揭示现有系统在领域特定错误和文化语义查询上的挑战。

Comments NOVAS Workshop at VLDB 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10724 2026-08-12 cs.CV 新提交 83%

InterPruner: Interactive Structured Pruning via Taylor-Implicit Criterion and Language-Prior Modulator for Multimodal Object Detection

InterPruner:面向多模态目标检测的、基于泰勒隐式准则与语言先验调制器的交互式结构化剪枝框架

Qi Ming, Zihan Yang, Shaoguang Huang, Si Sun, Hanqing Zhang, Nanqing Liu, Jiahui Lv, Juan Fang, Aleksandra Pizurica

机构 * Beijing University of Technology(北京工业大学) Southwest University(西南大学) China University of Geosciences Wuhan(中国地质大学(武汉)) Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学) Yunnan Normal University(云南师范大学) Beijing Forestry University(北京林业大学) Ghent University(根特大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出首个面向红外-可见光目标检测器的交互式结构化通道剪枝框架InterPruner,结合TIC、MIRA、SPCA方法,在FLIR数据集剪枝50%通道时mAP提升0.6%,代码将发布于GitHub。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07958 2026-08-11 cs.CV 新提交 83%

LIBAD: A Multimodal Anomaly Detection Benchmark for Li-Ion Battery Electrode Manufacturing

LIBAD:面向锂离子电池电极制造的多模态异常检测基准

Wenbo Sui, Daniel Lichau, Harold Phelippeau, Zhao Liu

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文推出首个面向锂离子电池电极制造的多模态异常检测基准LIBAD,提出DA-Core方法,其在核心集比例0.05时可降低误报率并缩短推理时间,优于相关基准方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05249 2026-08-10 cs.LG cs.AI 版本更新 83%

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis

PRISM:基于结构化多模态数据合成的感知优先级的评分标准内化

Xiaomin He, Dongling Xiao, Jiahao Xie, Ruiqi Lu, Qianle Wang, Zhongbin Guo, Wanxuan Sun

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract_cn);分类 cs.AI

AI总结 该研究针对多模态指令多要求重要性不等的问题,提出PRISM四阶段数据合成框架,结合PRISM-Eval评估,提升多模态大模型的多规则优先级感知指令遵循能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05817 2026-08-07 cs.CL 新提交 83%

M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

M$^3$R-Bench:一个用于基于证据的多模态隐喻理解的统一基准

Hong Jiang, Junnan Zhu, Jingwang Huang, Xiao Sun, Yuming Yang, Jiang Zhong, Ruirui Chen, Jingman Shi, Hao Wu, Nayu Liu, Xinyi Jiang, Kaiwen Wei

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出多模态隐喻理解基准M$^3$R-Bench,针对现有模型的跨模态证据-映射不匹配问题,构建M$^3$R-Reasoner方法,在多指标上超越GPT-5.5等现有模型。

Comments 6 figures and 5 tables. Hong Jiang, Junnan Zhu, and Jingwang Huang contributed equally. Jiang Zhong and Kaiwen Wei are corresponding authors. Code and data are available at https://github.com/hongshi4/M3R-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏