arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9111 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9111 篇

2201.02242 2022-01-10 eess.IV cs.CV 83%

A Keypoint Detection and Description Network Based on the Vessel Structure for Multi-Modal Retinal Image Registration

Aline Sindel, Bettina Hohberger, Sebastian Fassihi Dehcordi, Christian Mardin, Robert Lämmer, Andreas Maier, Vincent Christlein

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 6 pages, 4 figures, 1 table, accepted to BVM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.04873 2021-12-10 cs.CL 83%

Nice perfume. How long did you marinate in it? Multimodal Sarcasm Explanation

Poorav Desai, Tanmoy Chakraborty, Md Shad Akhtar

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted for publication in AAAI-2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.04264 2021-11-12 cs.CV 83%

Cross-Modal Object Tracking: Modality-Aware Representations and A Unified Benchmark

Chenglong Li, Tianhao Zhu, Lei Liu, Xiaonan Si, Zilin Fan, Sulan Zhai

专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments In Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.05971 2021-04-14 cs.CV 83%

Learning Multi-modal Information for Robust Light Field Depth Estimation

Yongri Piao, Xinxin Ji, Miao Zhang, Yukun Zhang

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.02636 2021-03-05 cs.CL cs.LG 83%

A Novel Context-Aware Multimodal Framework for Persian Sentiment Analysis

Kia Dashtipour, Mandar Gogate, Erik Cambria, Amir Hussain

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted in Neurocomputing

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.12126 2020-10-29 cs.CV 83%

Beyond the Deep Metric Learning: Enhance the Cross-Modal Matching with Adversarial Discriminative Domain Regularization

Li Ren, Kai Li, LiQiang Wang, Kien Hua

专题命中 多模态评测 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.08789 2020-07-16 cs.CV 83%

Borrow from Anywhere: Pseudo Multi-modal Object Detection in Thermal Imagery

Chaitanya Devaguptapu, Ninad Akolekar, Manuj M Sharma, Vineeth N Balasubramanian

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted at Perception Beyond Visible Spectrum Workshop, CVPR 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.07731 2020-05-29 eess.IV cs.CV 83%

Multi-modal Deep Guided Filtering for Comprehensible Medical Image Processing

Bernhard Stimpel, Christopher Syben, Franziska Schirrmacher, Philipp Hoelter, Arnd Dörfler, Andreas Maier

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Journal ref IEEE Transactions on Medical Imaging, vol. 39, no. 5, pp. 1703-1711, May 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.12238 2020-04-28 cs.CL 83%

MCQA: Multimodal Co-attention Based Network for Question Answering

Abhishek Kumar, Trisha Mittal, Dinesh Manocha

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.12676 2020-04-01 cs.CV 83%

xMUDA: Cross-Modal Unsupervised Domain Adaptation for 3D Semantic Segmentation

Maximilian Jaritz, Tuan-Hung Vu, Raoul de Charette, Émilie Wirbel, Patrick Pérez

专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted at CVPR 2020. For a demo video, see http://tiny.cc/xmuda

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.02649 2020-02-10 cs.CL 83%

Multimodal Matching Transformer for Live Commenting

Chaoqun Duan, Lei Cui, Shuming Ma, Furu Wei, Conghui Zhu, Tiejun Zhao

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.08948 2019-07-23 cs.CL 83%

Hindi Visual Genome: A Dataset for Multimodal English-to-Hindi Machine Translation

Shantipriya Parida, Ondřej Bojar, Satya Ranjan Dash

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.06322 2019-06-17 cs.CV cs.LG cs.RO 83%

Connecting Touch and Vision via Cross-Modal Prediction

Yunzhu Li, Jun-Yan Zhu, Russ Tedrake, Antonio Torralba

专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2019. Project Page: http://visgel.csail.mit.edu/

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.02858 2019-01-10 cs.CV 83%

Adaptive Feature Processing for Robust Human Activity Recognition on a Novel Multi-Modal Dataset

Mirco Moencks, Varuna De Silva, Jamie Roche, Ahmet Kondoz

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Working Draft

详情

展开后加载摘要…

URL PDF HTML 收藏
1602.08225 2016-02-29 cs.HC cs.CV cs.LG 83%

Multimodal Emotion Recognition Using Multimodal Deep Learning

Wei Liu, Wei-Long Zheng, Bao-Liang Lu

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05008 2026-06-04 cs.CV cs.AI cs.CL 83%

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks

M$^3$Eval: 通过认知基础视频任务的多模态记忆评估

Jie Huang, Ruixun Liu, Sirui Sun, Xinyi Yang, Yin Li, Yixin Zhu, Yiwu Zhong

机构 * School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Yuanpei College, Peking University(北京大学元培学院) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) School of Psychological and Cognitive Sciences, Peking University(北京大学心理学与认知科学学院) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出首个多模态模型记忆评估框架M$^3$Eval,通过认知心理学设计的视频任务系统评估模型在记忆保持、忠实性和鲁棒性上的表现,发现模型在并行视频流处理、干扰模式、时空记忆和符号记忆方面的显著缺陷。

Comments We present an evaluation designed for multi-modal memory in multi-modal models

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07632 2026-04-27 cs.AI cs.CL cs.CV cs.LG 83%

Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models

测试时匹配:在多模态模型中解锁组合推理

Yinglun Zhu, Jiancheng Zhang, Fuzhi Tang

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出测试时匹配算法,通过改进评估指标提升多模态模型的组合推理能力,使SigLIP-B16和GPT-4.1在Winoground等基准上取得新突破。

Comments To appear at ICLR 2026; extended results to generative multimodal models

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14707 2025-05-22 cs.MM cs.AI cs.CV 83%

CrypticBio: A Large Multimodal Dataset for Visually Confusing Biodiversity

Georgiana Manolache, Gerard Schouten, Joaquin Vanschoren

机构 * Fontys University of Applied Sciences(Fontys应用科学大学) Technical University of Eindhoven(埃因霍温技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments We present CrypticBio, the largest publicly available multimodal dataset of visually confusing species, specifically curated to support the development of AI models for biodiversity identification using images, language and spatiotemporal data

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04973 2024-07-09 cs.AI cs.CL cs.CV cs.LG 83%

LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Yijia Xiao, Edward Sun, Tianyu Liu, Wei Wang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments LogicVista benchmarks the logical reasoning of multimodal large language models in visual tasks

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12390 2024-07-04 cs.CV cs.AI cs.CL 83%

BLINK: Multimodal Large Language Models Can See but Not Perceive

Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A. Smith, Wei-Chiu Ma, Ranjay Krishna

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Multimodal Benchmark, Project Url: https://zeyofu.github.io/blink/, ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27378 2026-07-31 cs.CV cs.MM 新提交 82%

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

PanDent:面向牙科放射学中全面的牙级结构-语言一致性

Xiaohan Li, Xinyu Liu, Chang Liu, Sum Wing Au Yeung, Jun Liu, Yixuan Yuan, Hui Chen

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙医学院) Imperial College London(帝国理工学院) University of Science and Technology of China(中国科学技术大学) Department of Data and Systems Engineering, The University of Hong Kong(香港大学数据与系统工程系) Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系)

专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.MM

AI总结 本研究推出PanDent牙科OPG基准,经实验发现现有MLLM生成的牙科报告流畅但临床一致性差,在PanDent上微调可提升其结构-语言一致性,该基准可用于评估MLLM的牙级临床推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09654 2026-07-13 cs.CV cs.AI 新提交 82%

Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

十年视觉语言人工智能模型中准确性和视觉认知错误的演变

Shravan Murlidaran, Miguel P. Eckstein

机构 * Psychological & Brain Sciences, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校心理与脑科学系) Department of Computer Science, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校计算机科学系) Department of Electrical and Computer Engineering, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校电气与计算机工程系)

专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 研究十年间视觉语言模型进展,引入CSB数据集,评估模型在其上及MS-COCO样本中的准确性与视觉认知错误类型,发现MLLM消除简单与复杂场景描述准确性差距,几乎消除多数错误类型,为模型发展提供全面评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28215 2026-07-13 cs.AI cs.CL cs.LG cs.LO cs.MA 版本更新 82%

Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers

解释比单独预测更难:评估基于概念的MLLM解释作为ICL视觉分类器

Carmen Quiles-Ramírez, Leticia L. Rodríguez, Nicolás Martorell, Natalia Díaz-Rodríguez

专题命中 多模态评测 :MLLM(title_cn,abstract_cn);multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文通过五种形式化程度递增的条件,系统评估多模态大语言模型在少样本上下文学习中的基于概念的可解释性,发现解释比预测更难,且强制生成形式化解释会降低预测准确性。

Comments Accepted to the CompLearn Workshop at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07045 2026-05-15 cs.CV cs.AI 82%

VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing

VLRS-Bench: 一种面向遥感的视觉-语言推理基准

Zhiming Luo, Di Wang, Haonan Guo, Jing Zhang, Bo Du

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出VLRS-Bench,首个专注于复杂遥感推理的基准,包含2000个问题-答案对,涵盖14项任务和八个时间阶段,揭示现有MLLM在遥感任务中的瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08800 2026-05-12 cs.CV cs.AI 82%

PPU-Bench:Real World Benchmark for Personalized Partial Unlearning in Vision Language Models

PPU-Bench: 用于视觉语言模型个性化部分遗忘的现实世界基准

Jiahui Guang, Zexun Zhan, Zhenlin Xu, Cuiyun Gao, Haiyan Wang, Jing Li, Zhaoquan Gu, Yanchun Zhang

机构 * Harbin Institute of Technology(哈尔滨工业大学) Pengcheng Laboratory(鹏城实验室) The Hong Kong Polytechnic University(香港理工大学) Sichuan University(四川大学) Zhejiang Normal University(浙江师范大学)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出PPU-Bench,一个无需微调的现实世界基准,用于评估视觉语言模型中个性化部分遗忘的效果,通过24K多模态和单模态样本测试遗忘与保留的平衡及跨模态一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12436 2023-12-21 cs.CV cs.AI cs.CL cs.MM 82%

A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Chaoyou Fu, Renrui Zhang, Zihan Wang, Yubo Huang, Zhengye Zhang, Longtian Qiu, Gaoxiang Ye, Yunhang Shen, Mengdan Zhang, Peixian Chen, Sirui Zhao, Shaohui Lin, Deqiang Jiang, Di Yin, Peng Gao, Ke Li, Hongsheng Li, Xing Sun

专题命中 多模态评测 :multimodal(abstract,comments);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Total 120 pages. See our project at https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15006 2026-08-18 cs.CV cs.AI cs.MM 新提交 82%

MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Solving Geometry Problems

MetaReason:通过编辑元信息实现精确的交错多模态推理以解决几何问题

Penghao Yin, Haomin Wang, Qihong Tang, Xiaoye Qu, Hongjie Zhang, Xiao-Ping Zhang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本研究提出MetaReason框架,构建TutorGeo与ExamGeo数据集,结合监督微调与强化学习,实现几何问题的精确交错多模态推理,性能优于现有开源模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14767 2026-08-18 cs.CV cs.AI cs.CL cs.RO 新提交 82%

NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving

NARRATE:用于自动驾驶中以人为本解释的多模态真实世界澳大利亚驾驶数据集

Ashkan Yousefi Zadeh, Zishuo Zhu, Xiaomeng Li, Andry Rakotonirainy, Sebastien Glaser, Ronald Schroeter, Patricia Delhomme, Zahra Mehraban

机构 * ARC Training Centre for Automated Vehicles in Rural and Remote Regions (AVR3)(澳大利亚农村及偏远地区自动驾驶车辆ARC培训中心) Université Gustave Eiffel(Gustave Eiffel大学) Queensland University of Technology (QUT)(昆士兰科技大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究人员推出多模态真实世界澳大利亚驾驶数据集NARRATE,含2050个带注释事件,提供多类标签与情境意识注释,开展四项基准任务验证其价值,为开发以人为本的自动驾驶解释模型奠定基础。

Comments Accepted at The 19th European Conference on Computer Vision (ECCV 2026) DriveX Workshop (Foundation Models for Autonomous Driving)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13891 2026-08-17 cs.HC 新提交 82%

DepressionAgent: Reading, Listening, Seeing, and Deliberating Multimodal Evidence for Depression Risk Assessment

DepressionAgent:用于抑郁风险评估的阅读、倾听、观察与审议多模态证据框架

Fangjie Zhu, Haifeng Lu, Sicheng Zhao, Runhao Zeng, Xiping Hu

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

AI总结 研究针对现有多模态抑郁评估隐式特征融合的不足,提出以证据为中心的DepressionAgent框架,通过显式证据审议等机制在多基准上取得竞争力性能,且有效性与可检查性获多维度验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07867 2026-08-11 cs.LG 新提交 82%

CONFER: Conflict-Aware Evidence Negotiation for Regime-Calibrated Weak Supervision in Multimodal Emotion Recognition

CONFER:面向多模态情感识别中规则校准弱监督的冲突感知证据协商框架

Bojing Hou, Ruohao Li, Yitong Zhu, Luwen Yu, Yuyang Wang

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出CONFER框架,针对多模态情感识别中跨模态冲突与自我报告标签不可靠问题,通过图结构的冲突感知证据协商实现弱标签校准,在多数据集严格LOSO协议下取得具竞争力的情感识别准确率,提升了对弱标签损坏的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏