arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9111 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9111 篇

2505.20298 2026-01-27 cs.CL cs.AI cs.CV 82%

MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding

MangaVQA和MangaLMM:多模态漫画理解的基准和专用模型

Jeonghun Baek, Kazuki Egashira, Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Hikaru Ikuta, Kiyoharu Aizawa

机构 * The University of Tokyo(东京大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出MangaVQA和MangaLMM,用于多模态漫画理解的基准和专用模型,通过视觉问答和文本识别任务提升漫画叙事理解能力。

Comments EACL 2026 Findings. Project page: https://manga109.github.io/MangaVQA_LMM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13968 2026-01-27 cs.CV cs.AI cs.CL 82%

RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation

RotBench: 对多模态大语言模型识别图像旋转能力的评估

Tianyi Niu, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学夏洛特分校) Allen Institute for Artificial Intelligence(人工智能研究院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 RotBench评估了多模态大语言模型在识别图像旋转角度方面的性能,发现大多数模型难以区分90°和270°旋转,但能识别0°和180°图像,揭示了模型空间推理能力与人类的差距。

Comments EACL 2026 Camera-Ready. Code and data: https://github.com/tianyiniu/RotBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16520 2026-01-26 cs.CV cs.AI cs.CL 82%

TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning

TangramPuzzle: 通过组合空间推理评估多模态大语言模型

Daixian Liu, Jiayi Kuang, Yinghui Li, Yangning Li, Di Yin, Haoyu Cao, Xing Sun, Ying Shen, Hai-Tao Zheng, Liang Lin, Philip S. Yu

机构 * Tsinghua University(清华大学) Sun-Yat Sen University(孙逸人大学) Tencent Youtu Lab(腾讯优图实验室) University of Illinois Chicago(伊利诺伊大学香槟分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 TangramPuzzle通过几何基准测试评估多模态大语言模型的组合空间推理能力,发现模型在匹配轮廓时忽视几何约束,导致碎片变形。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17116 2026-01-16 cs.MM cs.AI cs.CV cs.IR 82%

The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding

CASTLE 2024数据集:推动多模态理解艺术的发展

Luca Rossetto, Werner Bailer, Duc-Tien Dang-Nguyen, Graham Healy, Björn Þór Jónsson, Onanong Kongmeesub, Hoang-Bao Le, Stevan Rudinac, Klaus Schöffmann, Florian Spiess, Allie Tran, Minh-Triet Tran, Quang-Linh Tran, Cathal Gurrin

机构 * Dublin City University(都柏林城市大学) JOANNEUM RESEARCH(JOANNEUM研究机构) University of Bergen(卑尔根大学) Reykjavik University(雷克雅未克大学) University of Amsterdam(阿姆斯特丹大学) Klagenfurt University(克雷克夫特大学) University of Basel(巴塞尔大学) VNU Ho Chi Minh University of Science(越南胡志明国家科学大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 CASTLE 2024数据集通过提供多视角视频和音频数据,推动多模态理解技术的发展。

Comments 7 pages, 6 figures, dataset available via https://castle-dataset.github.io/

Journal ref 2025 MM'25: Proceedings of the 33rd ACM International Conference on Multimedia (pp. 12629-12635)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06750 2026-01-13 cs.CV cs.AI cs.CL 82%

Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models

医疗多模态大语言模型的视点临床意图理解能力基准测试

Shaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng, Wenting Chen, Jie Liu, Linlin Shen

机构 * Shenzhen University(深圳大学) Stanford University(斯坦福大学) City University of Hong Kong(香港城市大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出MedGaze-Bench,首个评估医疗多模态大语言模型视点临床意图理解能力的基准测试,通过三维意图框架和陷阱QA机制,揭示现有模型在手术、急救和诊断任务中对意图理解的不足。

Comments 16 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02677 2026-01-07 cs.LG q-fin.RM q-fin.ST 82%

Uni-FinLLM: A Unified Multimodal Large Language Model with Modular Task Heads for Micro-Level Stock Prediction and Macro-Level Systemic Risk Assessment

Uni-FinLLM:一种具有模块化任务头的统一多模态大语言模型,用于微观层面的股票预测和宏观层面的系统性风险评估

Gongao Zhang, Haijiang Zeng, Lu Jiang

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

AI总结 Uni-FinLLM通过统一多模态大语言模型,结合模块化任务头,实现了对股票预测和系统性风险评估的高效联合建模,显著提升了预测准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20100 2026-01-07 cs.LG cs.AI cs.CL cs.CV 82%

MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations

MIRAGE:农业专家引导对话中多模态信息检索与推理的基准

Vardhan Dongre, Chi Gui, Shubham Garg, Hooshang Nayyeri, Gokhan Tur, Dilek Hakkani-Tür, Vikram S. Adve

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MIRAGE是一个用于农业专家引导对话中多模态信息检索与推理的基准,通过真实用户-专家交互数据,提供高保真的多模态推理评估平台。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13142 2026-01-01 cs.CV cs.CL cs.LG cs.MM cs.RO 82%

Holistic Evaluation of Multimodal LLMs on Spatial Intelligence

多模态大语言模型在空间智能方面的综合评估

Zhongang Cai, Yubo Wang, Qingping Sun, Ruisi Wang, Chenyang Gu, Wanqi Yin, Zhiqian Lin, Zhitao Yang, Chen Wei, Oscar Qian, Hui En Pang, Xuanke Shi, Kewang Deng, Xiaoyang Han, Zukai Chen, Jiaqi Li, Xiangyu Fan, Hanming Deng, Lewei Lu, Bo Li, Ziwei Liu, Quan Wang, Dahua Lin, Lei Yang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 本文提出EASI框架,用于评估多模态大语言模型在空间智能方面的综合表现,揭示GPT-5在SI中的优势与不足,并展示专有模型在困难任务上的劣势。

Comments Codebase: https://github.com/EvolvingLMMs-Lab/EASI/ ; Leaderboard: https://huggingface.co/spaces/lmms-lab-si/EASI-Leaderboard

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20938 2025-12-25 cs.HC 82%

Pioneering Multimodal Emotion Recognition in the Era of Large Models: From Closed Sets to Open Vocabularies

开创性多模态情感识别在大模型时代的探索:从封闭集到开放词汇

Jing Han, Zhiqiang Gao, Shihao Gao, Jialing Liu, Hongyu Chen, Zixing Zhang, Björn W. Schuller

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文首次系统评估了多模态大语言模型在开放词汇情感识别中的表现,发现两阶段三模态融合方法效果最佳,视频模态最为关键,同时揭示开放与封闭源LLM性能差距较小。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14620 2025-12-17 cs.CL cs.AI cs.CV 82%

JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction

JMMMU-Pro:通过Vibe基准构建的基于图像的日本多学科多模态理解基准

Atsuyuki Miyai, Shota Onohara, Jeonghun Baek, Kiyoharu Aizawa

机构 * The University of Tokyo(东京大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 JMMMU-Pro通过构建基于图像的多模态理解基准,评估LMMs在日文处理能力,提出Vibe基准构建方法以提高基准质量。

Comments Project page: https://mmmu-japanese-benchmark.github.io/JMMMU_Pro/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10701 2025-12-12 cs.LG 82%

HybridVFL: Disentangled Feature Learning for Edge-Enabled Vertical Federated Multimodal Classification

HybridVFL:面向边缘计算的垂直联邦多模态分类的解耦特征学习

Mostafa Anoosha, Zeinab Dehghani, Kuniko Paxton, Koorosh Aslansefat, Dhavalkumar Thakker

机构 * University of Hull(赫尔大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

AI总结 HybridVFL通过客户端特征解耦与服务器跨模态Transformer融合,提升边缘计算中多模态分类的隐私保护性能。

Comments 6 pages, 2 figures, 1 table. Accepted at UCC '25 (IEEE/ACM 18th International Conference on Utility and Cloud Computing), December 1-4, 2025, Nantes, France. DOI to be activated upon final publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08702 2025-12-10 cs.IR 82%

VI-MMRec: Similarity-Aware Training Cost-free Virtual User-Item Interactions for Multimodal Recommendation

VI-MMRec: 基于相似性感知的无成本虚拟用户-物品交互用于多模态推荐

Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li, Zitong Wan, Hewei Wang, Weijie Liu, Yijie Li, Edith C. H. Ngai

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

AI总结 VI-MMRec通过基于相似性感知的虚拟用户-物品交互,解决多模态推荐中的数据稀疏问题,提升模型性能且不增加训练成本。

Comments Accepted by KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17773 2025-12-05 cs.CV cs.AI cs.CL cs.LG 82%

KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models

KiVA:儿童启发的视觉类比用于测试大多模态模型

Eunice Yiu, Maan Qraitem, Anisa Noor Majhi, Charlie Wong, Yutong Bai, Shiry Ginosar, Alison Gopnik, Kate Saenko

机构 * University of California, Berkeley(加州大学伯克利分校) Boston University(波士顿大学) Google DeepMind(谷歌DeepMind) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 KiVA通过4300个日常物体视觉转换测试大模型的类比推理能力,发现儿童和成人表现优于现有模型,尤其在复杂任务上存在显著差距。

Comments 10 pages. Project website: https://ey242.github.io/kiva.github.io/. Benchmark and code: https://github.com/ey242/KiVA

Journal ref The Thirteenth International Conference on Learning Representations (ICLR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18980 2025-12-04 cs.CL cs.AI cs.CV 82%

IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web

IW-Bench: 评估大规模多模态模型的图像到网页转换能力

Hongcheng Guo, Wei Zhang, Junhao Chen, Yaonan Gu, Jian Yang, Junjia Du, Shaosheng Cao, Binyuan Hui, Tianyu Liu, Jianxin Ma, Chang Zhou, Zhoujun Li

机构 * Beihang University(北京航空航天大学) Alibaba Group(阿里巴巴集团) Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 IW-Bench提出了一种新的基准,用于评估大规模多模态模型在图像到网页转换中的性能,通过元素准确率和布局准确率以及五跳提示方法来提升评估效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23947 2025-11-19 cs.HC 82%

The Social Gaze of LLMs: A Literature Review of Multimodal Approaches to Human Behavior Understanding

Zihan Liu, Parisa Rabbani, Veda Duddu, Kyle Fan, Madison Lee, Yun Huang

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13248 2025-11-18 cs.CR 82%

DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents

Fuyao Zhang, Jiaming Zhang, Che Wang, Xiongtao Sun, Yurong Hao, Guowei Guan, Wenjie Li, Longtao Huang, Wei Yang Bryan Lim

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11503 2025-11-17 eess.SP 82%

SynthSoM-Twin: A Multi-Modal Sensing-Communication Digital-Twin Dataset for Sim2Real Transfer via Synesthesia of Machines

Junlong Chen, Ziwei Huang, Xuesong Cai, Xiang Cheng, Liuqing Yang

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10522 2025-11-13 cs.LG cs.AI cs.CV eess.AS 82%

Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction

Kaizhen Tan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15027 2025-11-10 cs.CL cs.AI cs.CV cs.HC 82%

InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback

Henry Hengyuan Zhao, Wenqi Pei, Yifei Tao, Haiyang Mei, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00020 2025-11-04 cs.AI cs.CL cs.CV 82%

Multimodal Detection of Fake Reviews using BERT and ResNet-50

Suhasnadh Reddy Veluru, Sai Teja Erukude, Viswa Chaitanya Marella

机构 * College of Business Administration Kansas State University Manhattan, USA(商学院学院 华盛顿州立大学 曼哈顿 美国) Department of Computer Science Kansas State University Manhattan, USA(计算机科学系 华盛顿州立大学 曼哈顿 美国)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published in IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00004 2025-11-04 cs.CY cs.AI cs.CL cs.CV 82%

Multimodal Learning with Augmentation Techniques for Natural Disaster Assessment

Adrian-Dinu Urse, Dumitru-Clementin Cercel, Florin Pop

机构 * Romanian Hub for Artificial Intelligence - HRIA(罗马尼亚人工智能枢纽 - HRIA)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at 2025 IEEE 21st International Conference on Intelligent Computer Communication and Processing (ICCP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25120 2025-10-30 cs.SI 82%

MMM-Fact: A Multimodal, Multi-Domain Fact-Checking Dataset with Multi-Level Retrieval Difficulty

Wenyan Xu, Dawei Xiang, Tianqi Ding, Weihai Lu

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

Comments Dataset link: https://huggingface.co/datasets/Wenyan0110/MMM-Fact

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22527 2025-10-28 astro-ph.IM astro-ph.GA cs.LG 82%

Multi-Modal Masked Autoencoders for Learning Image-Spectrum Associations for Galaxy Evolution and Cosmology

Morgan Himes, Samiksha Krishnamurthy, Andrew Lizarraga, Srinath Saikrishnan, Vikram Seenivasan, Jonathan Soriano, Ying Nian Wu, Tuan Do

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract)

Comments 8 pages, 3 figures, 1 table, accepted to NeurIPS 2025 Workshop ML4PS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21438 2025-10-27 cs.RO 82%

PREVENT: Proactive Risk Evaluation and Vigilant Execution of Tasks for Mobile Robotic Chemists using Multi-Modal Behavior Trees

Satheeshkumar Veeramani, Zhengxue Zhou, Francisco Munguia-Galeano, Hatem Fakhruldeen, Thomas Roddelkopf, Mohammed Faeik Ruzaij Al-Okby, Kerstin Thurow, Andrew Ian Cooper

机构 * Department of Chemistry and Material Innovation Factory, University of Liverpool(化学系和材料创新工厂,利物浦大学) Center for Life Science Automation, University of Rostock(生命科学自动化中心,罗斯托克大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

Comments 25 pages, 8 figures, paper submitted to Robotics and Autonomous Systems Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01243 2025-10-24 cs.CV cs.AI cs.CL 82%

Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants

Lixiong Qin, Shilong Ou, Miaoxuan Zhang, Jiangning Wei, Yuhang Zhang, Xiaoshuai Song, Yuchen Liu, Mei Wang, Weiran Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Normal University(北京师范大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 50 pages, 14 figures, 42 tables. NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09230 2025-10-13 cs.CV cs.AI cs.CL cs.LG 82%

Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras

Jindong Hong, Wencheng Zhang, Shiqin Qiao, Jianhai Chen, Jianing Qiu, Chuanyang Zheng, Qian Xu, Yun Ji, Qianyue Wen, Weiwei Sun, Hao Li, Huizhen Li, Huichao Wang, Kai Wu, Meng Li, Yijun He, Lingjie Luo, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) Peking University People’s Hospital(北京大学人民医院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15175 2025-10-13 cs.RO 82%

SHeRLoc: Synchronized Heterogeneous Radar Place Recognition for Cross-Modal Localization

Hanjun Kim, Minwoo Jung, Wooseong Yang, Ayoung Kim

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract)

Comments 9 pages, 9 figures, accepted to RA-L

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22385 2025-10-08 cs.CV cs.AI cs.CL 82%

Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment

Yue Zhang, Jilei Sun, Yunhui Guo, Vibhav Gogate

机构 * Department of Computer Science(计算机科学系) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11625 2025-10-07 cs.CL cs.AI cs.CV cs.LG 82%

MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering

Varun Srivastava, Fan Lei, Srija Mukhopadhyay, Vivek Gupta, Ross Maciejewski

机构 * School of Computing and Augmented Intelligence(计算与增强智能学院) Arizona State University(亚利桑那州立大学) Department of Computer Science(计算机科学系) International Institute of Information Technology(国际信息科技研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published as a conference paper at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00583 2025-10-02 cs.HC 82%

Rethinking Wine Tasting for Chinese Consumers: A Service Design Approach Enhanced by Multimodal Personalization

Xinyang Shan, Yuanyuan Xu, Tian Xia, Yinshan Lin

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏