arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9111 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9111 篇

2510.03878 2025-10-07 cs.CV cs.AI 81%

Multi-Modal Oral Cancer Detection Using Weighted Ensemble Convolutional Neural Networks

Ajo Babu George, Sreehari J R Ajo Babu George, Sreehari J R Ajo Babu George, Sreehari J R

机构 * Dicemed

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07640 2025-10-07 cs.MM cs.AI 81%

Synthesizing Sentiment-Controlled Feedback For Multimodal Text and Image Data

Puneet Kumar, Sarthak Malik, Balasubramanian Raman, Xiaobai Li

机构 * Center for Machine Vision and Signal Analysis, University of Oulu(机器视觉与信号分析中心,奥卢大学) Indian Institute of Technology Roorkee(印度理工学院罗尔基分校) State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01659 2025-10-03 cs.CL cs.AI 81%

MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization

Yinhong Liu, Jianfeng He, Hang Su, Ruixue Lian, Yi Nian, Jake Vincent, Srikanth Vishnubhotla, Robinson Piramuthu, Saab Mansour

机构 * AWS AI Labs(AWS人工智能实验室) Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01904 2025-10-03 cs.CV cs.AI 81%

What are You Looking at? Modality Contribution in Multimodal Medical Deep Learning

Christian Gapp, Elias Tappeiner, Martin Welk, Karl Fritscher, Elke Ruth Gizewski, Rainer Schubert

机构 * Institute of Biomedical Image Analysis UMIT TIROL -- Private University for Health Sciences(生物医学影像分析研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Contribution to Conference for Computer Assisted Radiology and Surgery (CARS 2025)

Journal ref Int J CARS (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24297 2025-10-01 cs.CL cs.AI 81%

Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs

Junying Wang, Zicheng Zhang, Ye Shen, Yalun Wu, Yingji Liang, Yijin Guo, Farong Wen, Wenzhe Li, Xuezhi Zhao, Qi Jia, Guangtao Zhai

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24888 2025-09-30 cs.CV cs.CL 81%

MMRQA: Signal-Enhanced Multimodal Large Language Models for MRI Quality Assessment

Fankai Jia, Daisong Gan, Zhe Zhang, Zhaochi Wen, Chenchen Dan, Dong Liang, Haifeng Wang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) ShanghaiTech University(上海理工大学) Southern University of Science and Technology(南方科技大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04740 2025-09-30 cs.CV cs.AI 81%

SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data

Samarth Mishra, Kate Saenko, Venkatesh Saligrama

机构 * Boston University(波士顿大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments ICCV 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23267 2025-09-30 cs.CV cs.AI cs.LG 81%

Learning Regional Monsoon Patterns with a Multimodal Attention U-Net

Swaib Ilias Mazumder, Manish Kumar, Aparajita Khan

机构 * 1Computer Science \& Engineering, Indian Institute of Technology Roorkee, India 2Computer Science \& Engineering, Indian Institute of Technology Ropar, India 3 Computer Science \& Engineering, Indian Institute of Technology (BHU) Varanasi, India

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in Geospatial AI and Applications with Foundation Models (GAIA) 2025, INSAIT and ELLIS Unit Sofia, Bulgaria

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23035 2025-09-30 cs.CV cs.AI 81%

Sensor-Adaptive Flood Mapping with Pre-trained Multi-Modal Transformers across SAR and Multispectral Modalities

Tomohiro Tanaka, Narumasa Tsutsumida

机构 * Graduate school of Science & Engineering, Saitama University, Japan(埼玉大学理工学部)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01565 2025-09-30 cs.CL cs.CV 81%

Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation

Li Zhou, Lutong Yu, Dongchu Xie, Shaohuan Cheng, Wenyan Li, Haizhou Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院) Chengdu Technological University(成都理工大学) University of Copenhagen(哥本哈根大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Cultural Analysis, Cultural Visual Understanding, Cultural Image Transcreation. Accepted by EMNLP 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20858 2025-09-26 cs.GR cs.CV cs.MM 81%

ArchGPT: Understanding the World's Architectures with Large Multimodal Models

Yuze Wang, Luo Yang, Junyi Wang, Yue Qi

机构 * State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室) School of Computer Science and Engineering(计算机科学与工程学院) Beihang University(北京航空航天大学) School of Computer Science and Technology(计算机科学与技术学院) Shandong University(山东大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00213 2025-09-26 cs.CV cs.AI 81%

Multimodal Deep Learning for Phyllodes Tumor Classification from Ultrasound and Clinical Data

Farhan Fuad Abir, Abigail Elliott Daly, Kyle Anderman, Tolga Ozmen, Laura J. Brattain

机构 * Department of Electrical and Computer Engineering, University of Central Florida(电子与计算机工程系,中央佛罗里达大学) Massachusetts General Hospital, Department of Surgery, Section of Breast Surgery(麻省总医院,外科部,乳腺外科) Department of Medicine, University of Central Florida College of Medicine(医学系,中央佛罗里达大学医学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments IEEE-EMBS International Conference on Body Sensor Networks (IEEE-EMBS BSN 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19952 2025-09-25 cs.CV cs.AI 81%

When Words Can't Capture It All: Towards Video-Based User Complaint Text Generation with Multimodal Video Complaint Dataset

Sarmistha Das, R E Zera Marveen Lyngkhoi, Kirtan Jain, Vinayak Goyal, Sriparna Saha, Manish Gupta

机构 * Indian Institute of Technology Patna(印度帕纳布理工大学) Microsoft, India(微软印度)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19274 2025-09-24 cs.CL cs.MM 81%

DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture

Arijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Nishant Mishra, Sriparna Saha

机构 * Indian Institute of Technology Patna(印度理工学院帕纳巴分校) Banasthali Vidyapeeth University(班纳萨利大学) Pandit Deendayal Energy University(德英德能源大学) Manipal University Jaipur(马哈拉施特拉邦大学贾伊普尔分校) Dwarkadas J. Sanghvi College of Engineering(德瓦尔卡斯J.桑格维工程学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments EMNLP MAINS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17740 2025-09-23 cs.CV cs.CL 81%

WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification

Yiwen Jiang, Deval Mehta, Siyuan Yan, Yaling Shen, Zimu Wang, Zongyuan Ge

机构 * Faculty of Engineering, Monash University(墨尔本大学工程学院) AIM for Health Lab, Faculty of IT, Monash University(墨尔本大学信息技术学院健康人工智能实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15701 2025-09-22 cs.CL cs.SD eess.AS 81%

Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment

Ke Wang, Wenning Wei, Yan Deng, Lei He, Sheng Zhao

机构 * Microsoft, Beijing, China(微软北京研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments submitted to ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10059 2025-09-15 cs.CV cs.AI 81%

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration

Yue Zhou, Litong Feng, Mengcheng Lan, Xue Yang, Qingyun Li, Yiping Ke, Xue Jiang, Wayne Zhang

机构 * Department of Physics, J.K. Institute of Science(J.K.科学研究院物理系) World Scientific University(世界科学大学) University of Intelligent Studies(智能研究大学) East China Normal University(华东师范大学) Nanyang Technological University(南洋理工大学) SenseTime Research(商汤科技研究院) Shanghai Jiao Tong University(上海交通大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 17 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09254 2025-09-12 cs.CV cs.MM 81%

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis

Jing Hao, Yuxuan Fan, Yanpeng Sun, Kaixin Guo, Lizhuo Lin, Jinrong Yang, Qi Yong H. Ai, Lun M. Wong, Hao Tang, Kuo Feng Hung

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院) The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学) National University of Singapore(新加坡国立大学) CVTE Sun Yat-sen University(孙中山大学) Department of Diagnostic Radiology, The University of Hong Kong(香港大学放射科) Imaging and Interventional Radiology, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院影像与介入放射科) School of Computer Science, Peking University(北京大学计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 40 pages, 26 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09014 2025-09-12 cs.CV cs.CL 81%

COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation

Umair Hassan

机构 * Independent Researcher(独立研究者)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 17 pages, 3 figures, 3 tables. Dataset available at https://huggingface.co/datasets/umairhassan02/urdu-translated-coco-captions-subset. Scripts and notebooks to reproduce results available at https://github.com/umair-hassan2/COCO-Urdu

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01306 2025-09-11 cs.AI cs.CV 81%

Multimodal Medical Disease Classification with LLaMA II

Christian Gapp, Elias Tappeiner, Martin Welk, Rainer Schubert

机构 * Institute of Biomedical Image Analysis UMIT TIROL(生物医学影像分析研究所 UMIT TIROL) UMIT TIROL – Private University for Health Sciences and Health Technology(UMIT TIROL – 健康科学与健康技术私立大学) VASCage – Centre on Clinical Stroke Research Innsbruck(VASCage – 神经科学临床研究中心 Innsbruck)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures, conference: AIRoV -- The First Austrian Symposium on AI, Robotics, and Vision 25.-27.3.2024, Innsbruck

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00785 2025-09-10 cs.AI cs.CV cs.LG 81%

GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning

Sahiti Yerramilli, Nilay Pande, Rynaa Grover, Jayant Sravan Tamarapalli

机构 * Google(谷歌) Waymo

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07050 2025-09-10 cs.CV cs.AI cs.CY 81%

Automated Evaluation of Gender Bias Across 13 Large Multimodal Models

Juan Manuel Contreras

机构 * Aymara

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06079 2025-09-09 cs.CL cs.CV 81%

Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge

Hao Liang, Ruitao Wu, Bohan Zeng, Junbo Niu, Wentao Zhang, Bin Dong

机构 * Peking University(北京大学) Beihang University(北航) Zhongguancun Academy(中关村学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00102 2025-09-09 cs.CV cs.CL cs.LG 81%

ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?

Pragati Shuddhodhan Meshram, Swetha Karthikeyan, Bhavya Bhavya, Suma Bhat

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05513 2025-09-09 cs.CV cs.AI cs.RO 81%

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

Ahad Jawaid, Yu Xiang

机构 * Department of Computer Science, The University of Texas at Dallas(德克萨斯大学达拉斯分校计算机科学系) Physical Automation, Inc.(Physical Automation 公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 4 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04469 2025-09-08 cs.CL cs.AI 81%

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing

David Berghaus, Armin Berger, Lars Hillebrand, Kostadin Cvejoski, Rafet Sifa

机构 * Fraunhofer IAIS(弗劳恩霍夫人工智能研究所) Lamarr Institute(拉马尔研究所)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08340 2025-09-01 cs.CV cs.AI 81%

Single Domain Generalization for Multimodal Cross-Cancer Prognosis via Dirac Rebalancer and Distribution Entanglement

Jia-Xuan Jiang, Jiashuai Liu, Hongtao Wu, Yifeng Wu, Zhong Wang, Qi Bi, Yefeng Zheng

机构 * Lanzhou University \& Westlake University Lanzhou China Xi'an Jiaotong University Xi'an China Westlake University \& Chinese University of Hong Kong University Hang Zhou, China Southern University of Science Lanzhou University Lanzhou China University of Amsterdam Amsterdam Netherland Westlake University Hangzhou China Lanzhou University \& Westlake University Xi'an Jiaotong University Westlake University \& Chinese University of Hong Kong University Lanzhou University University of Amsterdam Westlake University

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACMMM 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04836 2025-08-28 eess.IV cs.AI cs.CV 81%

PGAD: Prototype-Guided Adaptive Distillation for Multi-Modal Learning in AD Diagnosis

Yanfei Li, Teng Yin, Wenyi Shang, Jingyu Liu, Xi Wang, Kaiyang Zhao

机构 * Machine Intelligence Lab, College of Computer Science and Technology(机器智能实验室,计算机科学与技术学院) Sichuan University(四川大学) Department of Computer Science and Engineering(计算机科学与工程系) The Chinese University of Hong Kong(香港中文大学) Department of Neurosurgery(神经外科部门)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18740 2025-08-27 cs.CL cs.AI 81%

M3HG: Multimodal, Multi-scale, and Multi-type Node Heterogeneous Graph for Emotion Cause Triplet Extraction in Conversations

Qiao Liang, Ying Shen, Tiantian Chen, Lin Zhang

机构 * Tongji University(同济大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 16 pages, 8 figures. Accepted to Findings of ACL 2025

Journal ref Findings of ACL 2025 (2025) 11416-11431

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11169 2025-08-22 cs.CL cs.AI 81%

MuSeD: A Multimodal Spanish Dataset for Sexism Detection in Social Media Videos

Laura De Grazia, Pol Pastells, Mauro Vázquez Chas, Desmond Elliott, Danae Sánchez Villegas, Mireia Farrús, Mariona Taulé

机构 * University of Barcelona, CLiC-Language and Computing Center(巴塞罗那大学,CLiC语言与计算中心) University of Copenhagen, Department of Computer Science(哥本哈根大学,计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments COLM 2025 camera-ready version: expanded Section 4.3 with an additional experiment using an extended definition-based prompt (including a definition of sexist content), and applied minor corrections

详情

展开后加载摘要…

URL PDF HTML 收藏