arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9111 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9111 篇

2504.04974 2025-09-25 cs.CV cs.AI cs.CL cs.LG 82%

Towards Visual Text Grounding of Multimodal Large Language Model

Ming Li, Ruiyi Zhang, Jian Chen, Chenguang Wang, Jiuxiang Gu, Yufan Zhou, Franck Dernoncourt, Wanrong Zhu, Tianyi Zhou, Tong Sun

机构 * Adobe Research(Adobe研究院) University of Maryland(马里兰大学) University at Buffalo(布法罗大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15361 2025-09-22 cs.CL cs.AI cs.MM 82%

Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing

Zichen Wu, Hsiu-Yuan Huang, Yunfang Wu

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) MOE Key Laboratory of Computational Linguistics, Peking University(北京大学教育部计算语言学重点实验室) National Key Laboratory for Multimedia Information Processing, Peking University(北京大学国家多媒体信息处理重点实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15222 2025-09-19 cs.SD cs.CV cs.MM eess.AS eess.IV 82%

Two Web Toolkits for Multimodal Piano Performance Dataset Acquisition and Fingering Annotation

Junhyung Park, Yonghyun Kim, Joonhyung Bae, Kirak Kim, Taegyun Kwon, Alexander Lerch, Juhan Nam

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted to the Late-Breaking Demo Session of the 26th International Society for Music Information Retrieval (ISMIR) Conference, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12248 2025-09-18 cs.CV cs.AI cs.CL 82%

Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics

Yuriel Ryan, Rui Yang Tan, Kenny Tsu Wei Choo, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 27 pages, 8 figures, EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16633 2025-09-16 cs.CL cs.AI cs.MM 82%

GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View

Fenghua Cheng, Jinxiang Wang, Sen Wang, Zi Huang, Xue Li

机构 * The University of Queensland(昆士兰大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Updated version

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07030 2025-09-16 cs.CL cs.AI cs.CV cs.IR cs.LG 82%

FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering

Amirhossein Abaskohi, Spandana Gella, Giuseppe Carenini, Issam H. Laradji

机构 * Department of Computer Science(计算机科学系) The University of British Columbia(不列颠哥伦比亚大学) ServiceNow Research(ServiceNow研究)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07084 2025-09-12 cs.RO 82%

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models

Shucheng Huang, Freda Shi, Chen Sun, Jiaming Zhong, Minghao Ning, Yufeng Yang, Yukun Lu, Hong Wang, Amir Khajepour

机构 * MVSLab, Department of Mechanical and Mechatronics Engineering, University of Waterloo(滑铁卢大学机械与机电工程系MVSLab) CompLING Lab, David R. Cheriton School of Computer Science, University of Waterloo(滑铁卢大学大卫·R·切里顿计算机科学学院CompLING Lab) Department of Data and Systems Engineering, University of Hong Kong(香港大学数据与系统工程系) Department of Mechanical Engineering, University of New Brunswick(新不伦瑞克大学机械工程系) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动性学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract)

Comments This work has been accepted to IEEE Transactions on Vehicular Technology. Please refer to the copyright notice for additional information

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04658 2025-09-08 cs.RO 82%

Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision

Manish Kansana, Sindhuja Penchala, Shahram Rahimi, Noorbakhsh Amiri Golilarz

机构 * Department of Computer Science(计算机科学系) Engineering Mississippi State University Mississippi State, USA(工程学硕士州大学密西西比州) Department of Computer Science The University of Alabama Tuscaloosa, USA(计算机科学系阿拉巴马大学塔斯卡洛osa)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract)

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03986 2025-09-05 cs.CV cs.AI cs.CL cs.LG 82%

Promptception: How Sensitive Are Large Multimodal Models to Prompts?

Mohamed Insaf Ismithdeen, Muhammad Uzair Khattak, Salman Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学) Swiss Federal Institute of Technology Lausanne (EPFL)(洛桑联邦理工学院) Australian National University(澳大利亚国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03529 2025-09-05 cs.CL cs.AI eess.AS 82%

Multimodal Proposal for an AI-Based Tool to Increase Cross-Assessment of Messages

Alejandro Álvarez Castro, Joaquín Ordieres-Meré

机构 * AI master(人工智能硕士) Universidad Politécnica de Madrid(马德里理工大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments Presented at NLMLT2025 (https://airccse.org/csit/V15N16.html), 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07268 2025-09-03 cs.MM cs.CL cs.CV 82%

Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation

Jinyuan Li, Ziyan Li, Han Li, Jianfei Yu, Rui Xia, Di Sun, Gang Pan

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) NJUST(南京理工大学) College of Mathematics, Taiyuan University of Technology(数学学院,太原科技大学) Tianjin University of Science and Technology(天津科技大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Extension of our Findings of EMNLP 2023 & ACL 2024 paper, IEEE Transactions on Multimedia accepted on July 19, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19659 2025-08-28 cs.LG 82%

SCAR: A Characterization Scheme for Multi-Modal Dataset

Ri Su, Zhao Chen, Caleb Chen Cao, Nan Tang, Lei Chen

机构 * HKUST (GZ)(香港科技大学(珠海)) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16439 2025-08-28 cs.CY cs.AI cs.CL cs.GR cs.MM 82%

PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark

Adil Bahaj, Oumaima Fadi, Mohamed Chetouani, Mounir Ghogho

机构 * Mohammed 6 Polytechnic University(摩洛哥6号理工学院) International University of Rabat(拉巴特国际大学) Institut des Systèmes Intelligents et de Robotique(智能系统与机器人研究所)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15075 2025-08-26 cs.CL cs.AI cs.CV cs.LG 82%

Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs

Hao Wang, Pinzhi Huang, Jihan Yang, Saining Xie, Daisuke Kawahara

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments The first version of this paper mistakenly included a prompt injection phrase, which was inappropriate and unprofessional. Although we corrected the version on arXiv and withdrew from the conference, my co-authors and university strongly request a full withdrawal. Given the situation, I no longer have the authority to manage this paper, and withdrawing it from arXiv is the most responsible action

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13186 2025-08-20 cs.CL cs.AI cs.CV 82%

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

Shilong Li, Xingyuan Bu, Wenjie Wang, Jiaheng Liu, Jun Dong, Haoyang He, Hao Lu, Haozhe Zhang, Chenchen Jing, Zhen Li, Chuanhao Li, Jiayi Tian, Chenchen Zhang, Tianhao Peng, Yancheng He, Jihao Gu, Yuanxing Zhang, Jian Yang, Ge Zhang, Wenhao Huang, Wangchunshu Zhou, Zhaoxiang Zhang, Ruizhe Ding, Shilei Wen

机构 * Nanjing University(南京大学) Zhejiang University(浙江大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments The first two authors contribute equally, 26 pages, repo at https://github.com/MMBrowseComp/MM-BrowseComp

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09641 2025-08-14 cs.CE 82%

VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding

Zhaowei Liu, Xin Guo, Haotian Xia, Lingfeng Zeng, Fangqi Lou, Jinyi Niu, Mengping Li, Qi Qi, Jiahuan Li, Wei Zhang, Yinglong Wang, Weige Cai, Weining Shen, Liwen Zhang

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09864 2025-08-14 cs.LG cs.AI cs.CL cs.CV 82%

LUMA: A Benchmark Dataset for Learning from Uncertain and Multimodal Data

Grigor Bezirganyan, Sana Sellami, Laure Berti-Équille, Sébastien Fournier

机构 * Aix Marseille Univ, CNRS, LIS(阿维尼翁-马赛大学、国家科学研究中心、LIS)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments SIGIR 2025

Journal ref Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08093 2025-08-12 cs.CV cs.LG cs.MM eess.AS 82%

MDD-Net: Multimodal Depression Detection through Mutual Transformer

Md Rezwanul Haque, Md. Milon Islam, S M Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray

机构 * Centre for Pattern Analysis and Machine Intelligence, Department of Electrical and Computer Engineering, University of Waterloo(模式分析与机器智能中心,电气与计算机工程系,滑铁卢大学) School of Engineering and Computing, Department of Computer Science and Engineering, American University of Ras Al Khaimah(工程与计算学院,计算机科学与工程系,阿联酋拉线哈姆市美国大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted for the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17632 2025-08-12 cs.AI cs.CV cs.MM 82%

D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance

Renyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong Ng

机构 * School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学信息科学与技术学院(深圳校区)) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) College of Modern Engineering and the Engineering Research Center of Cyberspace, Yunnan University(云南大学现代工程学院及空天信息工程研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04350 2025-08-07 cs.CL cs.AI cs.CV cs.LG cs.MA 82%

Chain of Questions: Guiding Multimodal Curiosity in Language Models

Nima Iji, Kia Dashtipour

机构 * Edinburgh Napier University(爱丁堡纳皮尔大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04206 2025-08-07 cs.IR 82%

ViLLA-MMBench: A Unified Benchmark Suite for LLM-Augmented Multimodal Movie Recommendation

Fatemeh Nazary, Ali Tourani, Yashar Deldjoo, Tommaso Di Noia

专题命中 多模态评测 :multimodal(title,abstract);audio-visual(abstract)

Comments 17 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03733 2025-08-07 cs.LG cs.AI cs.CL cs.CV 82%

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning

Wenjie Li, Yujie Zhang, Haoran Sun, Yueqi Li, Fanrui Zhang, Mengzhe Xu, Victoria Borja Clausich, Sade Mellin, Renhao Yang, Chenrun Wang, Jethro Zih-Shuo Wang, Shiyi Yao, Gen Li, Yidong Xu, Hanyu Wang, Yilin Huang, Angela Lin Wang, Chen Shi, Yin Zhang, Jianan Guo, Luqi Yang, Renxuan Li, Yang Xu, Jiawei Liu, Yao Zhang, Lei Liu, Carlos Gutiérrez SanRomán, Lei Wang

机构 * College of Health Science and Technology, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院健康科学与技术学院) Shanghai Innovation Institute(上海创新研究院) Clinical Center for Sports Medicine, Department of Orthopaedics, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院骨科临床中心) School of Basic Medical Sciences, Intelligent Medicine Institute, Fudan University(复旦大学基础医学学院) Department of Hematology, The First Affiliated Hospital, College of Medicine, Zhejiang University(浙江大学医学院第一附属医院血液科) MoE Key Laboratory of Brain-Inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发智能感知与认知教育部重点实验室) Department of Public Health and Primary Care, University of Cambridge(剑桥大学公共卫生与初级保健学院) Department of Medicine, Faculty of Health Sciences, Universidad CEU Cardenal Herrera(CEU卡德纳尔-赫尔曼大学健康科学学院医学系) Faculty of Medicine, University of Helsinki(赫尔辛基大学医学院) X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院X-LANCE实验室) Department of Hepatobiliary Surgery, National Cancer Center / National Clinical Research Center for Cancer / Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(中国医学科学院肿瘤医院肝胆外科) Department of Surgery, The Ohio State University Wexner Medical Center, The James Comprehensive Cancer Center(俄亥俄州立大学韦克斯纳医学中心外科部,詹姆斯综合癌症中心) Ningbo Institute of Technology, Beihang University(北航宁波理工学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02745 2025-08-07 cs.CV cs.AI cs.CL 82%

AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity

Zhibin Lan, Liqiang Niu, Fandong Meng, Wenbo Li, Jie Zhou, Jinsong Su

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯人工智能研究院) Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15882 2025-08-06 cs.CV cs.AI cs.CL cs.LG 82%

Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark

Goeric Huybrechts, Srikanth Ronanki, Sai Muralidhar Jayanthi, Jack Fitzgerald, Srinivasan Veeravanallur

机构 * Amazon AGI(亚马逊人工智能研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10006 2025-08-01 cs.MM cs.AI cs.CV cs.LG 82%

HER2 Expression Prediction with Flexible Multi-Modal Inputs via Dynamic Bidirectional Reconstruction

Jie Qin, Wei Yang, Yan Su, Yiran Zhu, Weizhen Li, Yunyue Pan, Chengchang Pan, Honggang Qi

机构 * School of Computer Science and Technology, University of the Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院) School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院) Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎- Rocquencourt 国家信息与自动化所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒实验室)

专题命中 多模态评测 :multi-modal(title);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 8 pages,6 figures,3 tables,accepted by the 33rd ACM International Conference on Multimedia(ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01504 2025-07-16 cs.CV cs.AI cs.CL 82%

Following the Clues: Experiments on Person Re-ID using Cross-Modal Intelligence

Robert Aufschläger, Youssef Shoeb, Azarm Nowzad, Michael Heigl, Fabian Bally, Martin Schramm

机构 * Deggendorf Institute of Technology(德格多夫技术学院) Continental AG(大陆集团)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments accepted for publication at the 2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC 2025), taking place during November 18-21, 2025 in Gold Coast, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11341 2025-07-15 cs.MM cs.CL cs.CV cs.LG eess.IV 82%

MSVD-Indonesian: A Benchmark for Multimodal Video-Text Tasks in Indonesian

Willy Fitra Hendria

机构 * Independent Researcher(独立研究者)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments 10 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08238 2025-07-14 cs.LG 82%

Self-Supervised Learning-Based Multimodal Prediction on Prosocial Behavior Intentions

Abinay Reddy Naini, Zhaobo K. Zheng, Teruhisa Misu, Kumar Akash

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校) Honda Research Institute USA, Inc.(本田美国研究院)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract)

Comments 5 pages, 4 figures, published at ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04686 2025-07-08 cs.RO 82%

MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding

Jing Liang, Kasun Weerakoon, Daeun Song, Senthurbavan Kirubaharan, Xuesu Xiao, Dinesh Manocha

机构 * University of Maryland, College Park MD, 20740, USA(马里兰大学) Goerge Mason University, Fairfax, VA, 22030, USA(乔治·马歇尔大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22992 2025-07-01 cs.AI cs.CL cs.CV 82%

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning

Yulun Jiang, Yekun Chai, Maria Brbić, Michael Moor

机构 * EPFL(苏黎世联邦理工学院) ETH Zurich(苏黎世联邦理工学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏