arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9111 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9111 篇

2507.04270 2025-11-10 cs.CV cs.AI 81%

ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts

Sangbum Choi, Kyeongryeol Go, Taewoong Jang

机构 * Superb AI Seoul, South Korea(超霸AI首尔韩国)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02778 2025-11-05 cs.CV cs.CL 81%

VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation

Kevin Qinghong Lin, Yuhao Zheng, Hangyu Ran, Dantong Zhu, Dongxing Mao, Linjie Li, Philip Torr, Alex Jinpeng Wang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Project page: https://csu-jpg.github.io/VCode Github: https://github.com/CSU-JPG/VCode

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02495 2025-11-05 cs.CV cs.CL 81%

DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding

Zixuan Liu, Siavash H. Khajavi, Guangkai Jiang

机构 * Department of Computer Science(计算机科学系) Tulane University(Tulane 大学) Department of Industrial Engineering and Management(工业工程与管理系) Aalto University(Aalto 大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Advances in Neural Information Processing Systems 2025 (NeurIPS 2025), Poster, https://neurips.cc/virtual/2025/loc/san-diego/poster/121400

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27195 2025-11-05 cs.CV cs.CL cs.SI 81%

Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions

Caixin Kang, Yifei Huang, Liangyang Ouyang, Mingfang Zhang, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments ICCV2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13063 2025-11-04 cs.CV cs.CL 81%

PRISM2: Unlocking Multi-Modal General Pathology AI with Clinical Dialogue

Eugene Vorontsov, George Shaikovski, Adam Casson, Julian Viret, Eric Zimmermann, Neil Tenenholtz, Yi Kan Wang, Jan H. Bernhard, Ran A. Godrich, Juan A. Retamero, Jinru Shia, Mithat Gonen, Martin R. Weiser, David S. Klimstra, Razik Yousfi, Nicolo Fusi, Thomas J. Fuchs, Kristen Severson, Siqi Liu

机构 * Paige, NYC, NY United States(美国纽约市帕伊公司) Microsoft Research, Cambridge, MA United States(微软研究院) Memorial Sloan Kettering Cancer Center, NYC, NY United States(纪念斯隆凯特林癌症中心) University of Yale, New Haven, CT United States(耶鲁大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06594 2025-11-03 cs.CL cs.CV 81%

Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation

Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy F. Chen

机构 * IRIT, University of Toulouse, France(IRIT,图卢兹大学,法国) Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息通信研究所(I2R),A*STAR,新加坡) CNRS, IRIT, France(CNRS,IRIT,法国)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19028 2025-10-30 cs.CV cs.AI 81%

InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts

Tianchi Xie, Minzhi Lin, Mengchen Liu, Yilin Ye, Changjian Chen, Shixia Liu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22672 2025-10-29 cs.CV cs.CL cs.RO 81%

Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views

Anna Deichler, Jonas Beskow

机构 * KTH Royal Institute of Technology(皇家理工学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 10 pages, 6 figures, 2 tables. Accepted to the NeurIPS 2025 Workshop on SPACE in Vision, Language, and Embodied AI (SpaVLE). Dataset: https://huggingface.co/datasets/annadeichler/KTH-ARIA-referential

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22622 2025-10-28 cs.CR cs.CV cs.MM 81%

DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection

Kangran Zhao, Yupeng Chen, Xiaoyu Zhang, Yize Chen, Weinan Guan, Baicheng Chen, Chengzhe Sun, Soumyya Kanti Datta, Qingshan Liu, Siwei Lyu, Baoyuan Wu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University at Buffalo, State University of New York(纽约州立大学布法罗分校) Nanjing University of Posts and Telecommunications(南京邮电大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20759 2025-10-28 cs.CV cs.AI 81%

PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding

Ansel Blume, Jeonghwan Kim, Hyeonjeong Ha, Elen Chatikyan, Xiaomeng Jin, Khanh Duy Nguyen, Nanyun Peng, Kai-Wei Chang, Derek Hoiem, Heng Ji

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California Los Angeles(加州大学洛杉矶分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025 Spotlight; project page: https://wjdghks950.github.io/partonomy.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20381 2025-10-24 cs.CL cs.AI 81%

VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation

Son T. Luu, Trung Vo, Hiep Nguyen, Khanh Quoc Tran, Kiet Van Nguyen, Vu Tran, Ngan Luu-Thuy Nguyen, Le-Minh Nguyen

机构 * Japan Advanced Institute of Science and Technology(日本先进科学研究院) University of Information Technology(信息技术大学) Vietnam National University(越南国家大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments VLSP 2025 MLQA-TSR Share Task

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19892 2025-10-24 cs.CL cs.AI 81%

Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities

Nishant Balepur, Dang Nguyen, Dayeon Ki

机构 * University of Maryland(马里兰大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted as a Spotlight paper at the EMNLP 2025 Wordplay Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19451 2025-10-23 cs.CV cs.MM 81%

Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis

Xueqi Ma, Yanbei Jiang, Sarah Erfani, James Bailey, Weifeng Liu, Krista A. Ehinger, Jey Han Lau

机构 * The University of Melbourne(墨尔本大学) China University of Petroleum (East China)(中国石油大学(华东))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19183 2025-10-23 cs.CV cs.AI 81%

PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning

Fengyuan Sun, Hui Chen, Xinhao Xu, Dandan Zheng, Jingdong Chen, Jun Zhou, Jungong Han, Guiguang Ding

机构 * School of Software, Tsinghua University(清华大学软件学院) Ant Group(蚂蚁集团) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00711 2025-10-23 cs.LG cs.AI cs.CV 81%

QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training

Wei Dai, Peilin Chen, Chanakya Ekbote, Paul Pu Liang

机构 * MIT Media Lab(MIT媒体实验室) MIT EECS(MIT电子工程与计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted as Oral at NeurIPS 2025. Revision after camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18014 2025-10-22 cs.CV cs.MM 81%

ManzaiSet: A Multimodal Dataset of Viewer Responses to Japanese Manzai Comedy

Kazuki Kawamura, Kengo Nakai, Jun Rekimoto

机构 * Sony CSL Kyoto(索尼 CSL 京都) The University of Tokyo(东京大学) Yoshimoto Kogyo Holdings Co., Ltd.(吉村工业株式会社)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments ICCV 2025 Workshop on Affective & Behavior Analysis in-the-Wild (ABAW), Honolulu, HI, USA (Oct 19, 2025, HST). 11 pages, 5 figures

Journal ref ICCV 2025 Workshops (ICCVW) / CVF Open Access

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02537 2025-10-22 cs.CV cs.AI 81%

VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning

Hao Yan, Xingchen Liu, Hao Wang, Zhenbiao Cao, Handong Zheng, Liang Yin, Xinxing Su, Zihao Chen, Jihao Wu, Minghui Liao, Chao Weng, Wei Chen, Yuliang Liu, Xiang Bai

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18601 2025-10-21 cs.CL cs.AI 81%

Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators

Jongwoo Ko, Sungnyun Kim, Sungwoo Cho, Se-Young Yun

机构 * Microsoft(微软公司) KAIST AI(韩国科学技术院人工智能研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17455 2025-10-21 cs.CL cs.AI 81%

Towards Evaluating Proactive Risk Awareness of Multimodal Language Models

Youliang Yuan, Wenxiang Jiao, Yuejin Xie, Chihao Shen, Menghan Tian, Wenxuan Wang, Jen-tse Huang, Pinjia He

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院) Xiaohongshu Inc.(小红书公司) Renmin University of China(中国人民大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2025 (Track on Datasets and Benchmarks)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15684 2025-10-20 cs.CV cs.AI 81%

Towards Label-Free Brain Tumor Segmentation: Unsupervised Learning with Multimodal MRI

Gerard Comas-Quiles, Carles Garcia-Cabrera, Julia Dietlmeier, Noel E. O'Connor, Ferran Marques

机构 * Universitat Politècnica de Catalunya (UPC)(西班牙巴塞罗那理工大学) University College Dublin (UCD)(都柏林大学) Dublin City University (DCU)(都柏林城市大学) Insight Research Ireland Center For Data Analytics(爱尔兰数据分析研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, 5 figures, BraTS GoAT 2025 challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08636 2025-10-20 cs.CV cs.AI 81%

Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models

Xingrui Wang, Wufei Ma, Tiezheng Zhang, Celso M de Melo, Jieneng Chen, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published in CVPR 2025 as Highlight. Data and code are released at https://github.com/XingruiWang/Spatial457

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14958 2025-10-17 cs.CV cs.CL 81%

MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning

Weikang Shi, Aldrich Yu, Rongyao Fang, Houxing Ren, Ke Wang, Aojun Zhou, Changyao Tian, Xinyu Fu, Yuxuan Hu, Zimu Lu, Linjiang Huang, Si Liu, Rui Liu, Hongsheng Li

机构 * Multimedia Laboratory (MMLab), The Chinese University of Hong Kong(中文大学多媒体实验室) Huawei Research(华为研究) BUAA(北京航空学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Project Page: https://mathcanvas.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14307 2025-10-17 cs.CL cs.AI 81%

MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking

Sathyanarayanan Ramamoorthy, Vishwa Shah, Simran Khanuja, Zaid Sheikh, Shan Jie, Ann Chia, Shearman Chua, Graham Neubig

机构 * Carnegie Mellon University(卡内基梅隆大学) Defence Science and Technology Agency(国防科学与技术局)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19330 2025-10-16 eess.SP cs.AI cs.HC cs.LG cs.MM 81%

LibEMER: A novel benchmark and algorithms library for EEG-based Multimodal Emotion Recognition

Zejun Liu, Yunshan Chen, Chengxi Xie, Yugui Xie, Huan Liu

机构 * XJTU-POLIMI Joint School, Xi'an Jiaotong University School of Computer Science Technology, Xi'an Jiaotong University MIGU Video Co., Ltd., Shanghai, China

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10412 2025-10-16 cs.LG cs.AI cs.CL 81%

Time-IMM: A Dataset and Benchmark for Irregular Multimodal Multivariate Time Series

Ching Chang, Jeehyun Hwang, Yidan Shi, Haixin Wang, Wen-Chih Peng, Tien-Fu Chen, Wei Wang

机构 * University of California, Los Angeles(加州大学洛杉矶分校) National Yang Ming Chiao Tung University(国立阳明交通大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted by the NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13211 2025-10-16 cs.CV cs.AI 81%

MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance

Subin Kim, Hoonrae Kim, Jihyun Lee, Yejin Jeon, Gary Geunbae Lee

机构 * KT Corporation, Republic of Korea(韩国KT公司) Graduate School of Artificial Intelligence, POSTECH, Republic of Korea(POSTECH人工智能研究生院) Computer Science and Engineering, POSTECH, Republic of Korea(POSTECH计算机科学与工程系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10965 2025-10-14 cs.CL cs.AI 81%

Judge Before Answer: Can MLLM Discern the False Premise in Question?

Jidong Li, Lingyong Fang, Haodong Zhao, Sufeng Duan, Gongshen Liu

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) Inner Mongolia Research Institute, Shanghai Jiao Tong University(上海交通大学内蒙古研究院)

专题命中 多模态评测 :MLLM(title);multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10546 2025-10-14 cs.CV cs.AI 81%

GLOFNet -- A Multimodal Dataset for GLOF Monitoring and Prediction

Zuha Fatima, Muhammad Anser Sohaib, Muhammad Talha, Sidra Sultana, Ayesha Kanwal, Nazia Perwaiz

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08964 2025-10-13 cs.CV cs.CL 81%

Unleashing Perception-Time Scaling to Multimodal Reasoning Models

Yifan Li, Zhenghao Chen, Ziheng Wu, Kun Zhou, Ruipu Luo, Can Zhang, Zhentao He, Yufei Zhan, Wayne Xin Zhao, Minghui Qiu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理重点实验室) ByteDance(字节跳动) University of California, San Diego(加州大学圣地亚哥分校) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06019 2025-10-10 cs.CV cs.AI eess.IV eess.SP 81%

BRIGHT: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response

Hongruixuan Chen, Jian Song, Olivier Dietrich, Clifford Broni-Bediako, Weihao Xuan, Junjue Wang, Xinlei Shao, Yimin Wei, Junshi Xia, Cuiling Lan, Konrad Schindler, Naoto Yokoya

机构 * Graduate School of Frontier Sciences, The University of Tokyo(东京大学前沿科学研究生院) RIKEN Center for Advanced Intelligence Project (AIP), RIKEN(日本理化学研究院先进智能项目中心) Department of Photogrammetry and Remote Sensing, ETH Zürich(苏黎世联邦理工学院测绘与遥感系) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏