arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9087 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9087 篇

2509.09160 2025-09-12 cs.CL cs.AI 84%

Target-oriented Multimodal Sentiment Classification with Counterfactual-enhanced Debiasing

Zhiyue Liu, Fanrong Ma, Xin Ling

机构 * School of Computer, Electronics and Information(计算机、电子与信息学院) Guangxi University(广西大学) Guangxi Key Laboratory of Multimedia Communications(广西多媒体通信与网络技术重点实验室) School of Sociology and Anthropology(社会学与人类学学院) Sun Yat-sen University(中山大学)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments Accepted by the IEEE International Conference on Multimedia and Expo (ICME 2025). © 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08008 2025-09-11 cs.SI cs.AI cs.MM 84%

A New Dataset and Benchmark for Grounding Multimodal Misinformation

Bingjian Yang, Danni Xu, Kaipeng Niu, Wenxuan Liu, Zheng Wang, Mohan Kankanhalli

机构 * National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(多媒体软件国家工程研究中心,计算机科学学院,武汉大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学) School of Computer Science, Peking University(计算机科学学院,北京大学) State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室,北京大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

Comments 6 pages, 5 figures, ACM Multimedia 2025 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20670 2025-09-10 cs.CV cs.MM 84%

"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection

Anastasios Skoularikis, Stefanos-Iordanis Papadopoulos, Symeon Papadopoulos, Panagiotis C. Petrantonakis

机构 * Department of Electrical & Computer Engineering, Aristotle University of Thessaloniki(电气与计算机工程系,阿基米德大学塞萨洛尼基分校) Information Technology Institute, Centre for Research & Technology Hellas(信息科技研究所,希腊研究中心)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13111 2025-09-09 cs.CV cs.CL cs.LG 84%

MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs

Erik Daxberger, Nina Wenzel, David Griffiths, Haiming Gang, Justin Lazarow, Gefen Kohavi, Kai Kang, Marcin Eichner, Yinfei Yang, Afshin Dehghan, Peter Grasch

机构 * Apple(苹果公司)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04844 2025-09-08 cs.MM cs.AI cs.IR 84%

REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-Experts

Xinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang, Hongbo Xu, Hao Xu, Yubin Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00484 2025-09-03 cs.CV cs.AI 84%

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding

Zhihong Zhang, Xiaojian Huang, Jin Xu, Zhuodong Luo, Xinzhi Wang, Jiansheng Wei, Xuejin Chen

机构 * University of Science and Technology of China(中国科学技术大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments https://videorewardbench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12263 2025-09-01 cs.CV cs.AI 84%

Region-Level Context-Aware Multimodal Understanding

Hongliang Wei, Xianqi Zhang, Xingtao Wang, Xiaopeng Fan, Debin Zhao

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院长) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20068 2025-08-28 cs.CL cs.CV cs.LG 84%

11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis

Chengzu Li, Wenshan Wu, Huanyu Zhang, Qingtao Li, Zeyu Gao, Yan Xia, José Hernández-Orallo, Ivan Vulić, Furu Wei

机构 * Microsoft Research(微软研究院) Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) Department of Oncology, University of Cambridge(癌症部门,剑桥大学) Leverhulme Centre for the Future of Intelligence, University of Cambridge(未来智能中心,剑桥大学) VRAIN, Universitat Politècnica de València(VRAIN,巴塞罗那理工大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments 9 pages, 4 figures (22 pages, 7 figures, 7 tables including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01490 2025-08-28 q-bio.GN cs.AI cs.CV cs.LG q-bio.TO stat.AP 84%

A Large-Scale Benchmark of Cross-Modal Learning for Histology and Gene Expression in Spatial Transcriptomics

Rushin H. Gindra, Giovanni Palla, Mathias Nguyen, Sophia J. Wagner, Manuel Tran, Fabian J Theis, Dieter Saur, Lorin Crawford, Tingying Peng

机构 * Helmholtz Munich, Germany(海德堡-慕尼黑赫尔姆霍尔茨研究中心) Technical University Munich, Germany(慕尼黑技术大学) Microsoft Research, USA(微软研究院)

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments The code is accessible at: https://github.com/peng-lab/hescape

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15370 2025-08-22 cs.CL cs.AI 84%

Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation

Yichi Zhang, Yao Huang, Yifan Wang, Yitong Sun, Chang Liu, Zhe Zhao, Zhengwei Fang, Huanran Chen, Xiao Yang, Xingxing Wei, Hang Su, Yinpeng Dong, Jun Zhu

机构 * Department of Computer Science and Technology, College of AI, Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(计算机科学与技术系、人工智能学院、人工智能研究所、清华-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学) Institute of Artificial Intelligence, Beihang University(人工智能研究院、北航) RealAI

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments For Appendix, please refer to arXiv:2406.07057

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13796 2025-08-20 cs.CV cs.AI 84%

A Fully Transformer Based Multimodal Framework for Explainable Cancer Image Segmentation Using Radiology Reports

Enobong Adahada, Isabel Sassoon, Kate Hone, Yongmin Li

机构 * Department of Computer Science(计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17303 2025-08-20 eess.IV cs.AI cs.CV 84%

A Versatile Pathology Co-pilot via Reasoning Enhanced Multimodal Large Language Model

Zhe Xu, Ziyi Liu, Junlin Hou, Jiabo Ma, Cheng Jin, Yihui Wang, Zhixuan Chen, Zhengyu Zhang, Fuxiang Huang, Zhengrui Guo, Fengtao Zhou, Yingxue Xu, Xi Wang, Ronald Cheong Kin Chan, Li Liang, Hao Chen

机构 * Hong Kong University of Science and Technology(香港科技大学) Southern Medical University(南方医科大学) Nanfang Hospital and School of Basic Medical Sciences(南方医院和基础医学科学学院) Chinese University of Hong Kong(香港中文大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07766 2025-08-12 cs.CV cs.AI 84%

UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models

Jinke Li, Jiarui Yu, Chenxing Wei, Hande Dong, Qiang Lin, Liangjing Yang, Zhicai Wang, Yanbin Hao

机构 * Zhejiang University(浙江大学) Tencent(腾讯) Shenzhen University(深圳大学) Hefei University of Technology(合肥工业大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted at ACM MM 2025 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01274 2025-08-05 cs.AI cs.CL 84%

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan

Jui-Ming Yao, Bing-Cheng Xie, Sheng-Wei Peng, Hao-Yuan Chen, He-Rong Zheng, Bing-Jia Tan, Peter Shaojui Wang, Shun-Feng Su

机构 * National Taiwan University of Science and Technology(台湾科技大学) University of London(伦敦大学) National Taiwan University(台湾大学)

专题命中 多模态评测 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22367 2025-07-31 cs.CL cs.MM 84%

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors

Jia Li, Yichao He, Jiacheng Xu, Tianhao Luo, Zhenzhen Hu, Richang Hong, Meng Wang

机构 * Hefei University of Technology(合肥工业大学)

专题命中 多模态评测 :multimodal(title);cross-modal(abstract);audio-visual(abstract);分类 cs.CL、cs.MM

Comments 8 pages, 3 figures, ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20217 2025-07-30 cs.RO cs.AI cs.CV 84%

Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots

Wei Cui, Haoyu Wang, Wenkang Qin, Yijie Guo, Gang Han, Wen Zhao, Jiahang Cao, Zhang Zhang, Jiaru Zhong, Jingkai Sun, Pihai Sun, Shuai Shi, Botuo Jiang, Jiahao Ma, Jiaxu Wang, Hao Cheng, Zhichao Liu, Yang Wang, Zheng Zhu, Guan Huang, Jian Tang, Qiang Zhang

机构 * X-Humanoid GigaAI Project(GigaAI项目)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Tech Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17776 2025-07-29 cs.CV cs.MM 84%

Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search

Shuyu Yang, Yaxiong Wang, Li Zhu, Zhedong Zheng

机构 * Xi’an Jiaotong University(西安交通大学) Hefei University of Technology(合肥工业大学) University of Macau(澳门大学)

专题命中 多模态评测 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03328 2025-07-29 cs.CV cs.AI cs.NE 84%

Visual Enumeration Remains Challenging for Multimodal Generative AI

Alberto Testolin, Kuinan Hou, Marco Zorzi

机构 * Department of General Psychology and Department of Mathematics University of Padova(帕多瓦大学心理学系和数学系) Department of General Psychology University of Padova(帕多瓦大学心理学系) Department of General Psychology and Padova Neuroscience Center University of Padova(帕多瓦大学心理学系和帕多瓦神经科学中心) IRCSS San Camillo Hospital, Venice-Lido(威尼斯利多医院IRCSS桑卡莫医院)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10510 2025-07-25 cs.CV cs.CL 84%

DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts

Tobias Braun, Mark Rothermel, Marcus Rohrbach, Anna Rohrbach

机构 * Technical University of Darmstadt \& hessian.AI, Germany

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments ICML 2025 version. 9 pages main paper, 35 pages with appendix, 18 figures and 7 tables. Corrected two inconsistent numbers in Table 2

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14544 2025-07-22 cs.CV cs.AI 84%

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025

Sujata Gaihre, Amir Thapa Magar, Prasuna Pokharel, Laxmi Tiwari

机构 * NCIT(尼泊尔信息技术研究所) Fusemachine(Fusemachine公司) Logictronix Technologies(Logictronix Technologies公司)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

Comments accepted to ImageCLEF 2025, to be published in the lab proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06852 2025-07-17 cs.CV cs.AI 84%

Position Prediction Self-Supervised Learning for Multimodal Satellite Imagery Semantic Segmentation

John Waithaka, Moise Busogi

机构 * Carnegie Mellon University Africa(卡内基梅隆大学非洲分校)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15426 2025-07-17 cs.CV cs.AI 84%

Visual Position Prompt for MLLM based Visual Grounding

Wei Tang, Yanpeng Sun, Qinying Gu, Zechao Li

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06110 2025-07-16 cs.CL cs.AI 84%

Multimodal Sentiment Analysis on CMU-MOSEI Dataset using Transformer-based Models

Jugal Gajjar, Kaustik Ranaware

机构 * George Washington University(乔治·华盛顿大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05007 2025-07-11 cs.CV cs.AI 84%

Multi-modal Representations for Fine-grained Multi-label Critical View of Safety Recognition

Britty Baby, Vinkle Srivastav, Pooja P. Jain, Kun Yuan, Pietro Mascagni, Nicolas Padoy

机构 * University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France(斯特拉斯堡大学,法国国家科学研究中心,法国国家卫生研究院,ICube,UMR7357,斯特拉斯堡,法国) Fondazione Policlinico Universitario A. Gemelli IRCCS, Università Cattolica del Sacro Cuore, Rome, Italy(A. Gemelli IRCCS大学医院,罗马,意大利) CAMP, Technische Universität München, Munich, Germany(慕尼黑技术大学,德国) Institute of Image-Guided Surgery, IHU Strasbourg, Strasbourg, France(影像引导手术研究所,斯特拉斯堡,法国)

专题命中 多模态评测 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04333 2025-07-08 cs.CV cs.CL 84%

Computed Tomography Visual Question Answering with Cross-modal Feature Graphing

Yuanhe Tian, Chen Su, Junwen Duan, Yan Song

机构 * University of Washington(华盛顿大学) University of Science and Technology of China(中国科学技术大学) Central South University(中南大学)

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21604 2025-06-30 cs.IR cs.AI cs.CV cs.HC cs.LG 84%

Evaluating VisualRAG: Quantifying Cross-Modal Performance in Enterprise Document Understanding

Varun Mannam, Fang Wang, Xin Chen

机构 * Amazon(亚马逊)

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Conference: KDD conference workshop: https://kdd-eval-workshop.github.io/genai-evaluation-kdd2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21277 2025-06-27 cs.CV cs.CL 84%

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Qize Yang, Shimin Yao, Weixuan Chen, Shenghao Fu, Detao Bai, Jiaxing Zhao, Boyuan Sun, Bowen Yin, Xihan Wei, Jingren Zhou

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

专题命中 多模态评测 :omni-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19469 2025-06-25 cs.CV cs.AI 84%

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning

Pengfei Hao, Shuaibo Li, Hongqiu Wang, Zhizhuo Kou, Junhang Zhang, Guang Yang, Lei Zhu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Hong Kong University of Science and Technology(香港科技大学) Department of Thoracic Surgery, the Seventh Affiliated Hospital, Sun Yat-sen University(中山大学第七附属医院胸外科部门) Bioengineering/Imperial-X, Imperial College London(生物工程/Imperial-X,帝国理工学院伦敦分校) ROAS Thrust, Hong Kong University of Science and Technology (Guangzhou)(ROAS项目,香港科技大学(广州)) Department of Electronic and Computer Engineering, Hong Kong SAR(香港特别行政区电子与计算机工程系)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18512 2025-06-25 eess.IV cs.CL cs.CV q-bio.QM 84%

MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis

Yuting Zhang, Kaishen Yuan, Hao Lu, Yutao Yue, Jintai Chen, Kaishun Wu

机构 * The Hong Kong University of Science & Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01173 2025-06-24 cs.AI cs.CV cs.LG 84%

Reasoning Limitations of Multimodal Large Language Models. A Case Study of Bongard Problems

Mikołaj Małkiński, Szymon Pawlonka, Jacek Mańdziuk

机构 * Warsaw University of Technology, Warsaw, Poland(华沙技术大学) AGH University of Krakow, Krakow, Poland(克拉科夫AGH大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to The Forty-Second International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏