arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9119 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9119 篇

1904.01356 2019-04-03 cs.CV cs.CL 81%

Aiding Intra-Text Representations with Visual Context for Multimodal Named Entity Recognition

Omer Arshad, Ignazio Gallo, Shah Nawaz, Alessandro Calefati

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.07427 2019-02-13 cs.CL cs.CV cs.IR 81%

Multimodal Sentiment Analysis: Addressing Key Issues and Setting up the Baselines

Soujanya Poria, Navonil Majumder, Devamanyu Hazarika, Erik Cambria, Alexander Gelbukh, Amir Hussain

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments IEEE Intelligence Systems. arXiv admin note: substantial text overlap with arXiv:1707.09538

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.12276 2018-12-05 cs.CL cs.AI cs.LG 81%

Improving Hospital Mortality Prediction with Medical Named Entities and Multimodal Learning

Mengqi Jin, Mohammad Taha Bahadori, Aaron Colak, Parminder Bhatia, Busra Celikkaya, Ram Bhakta, Selvan Senthivel, Mohammed Khalilia, Daniel Navarro, Borui Zhang, Tiberiu Doman, Arun Ravi, Matthieu Liger, Taha Kass-hout

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Machine Learning for Health (ML4H) Workshop at NeurIPS 2018 arXiv:1811.07216

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.03695 2018-09-12 cs.CL cs.AI 81%

Evaluating Multimodal Representations on Sentence Similarity: vSTS, Visual Semantic Textual Similarity Dataset

Oier Lopez de Lacalle, Aitor Soroa, Eneko Agirre

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Journal ref ICCV17: second workshop on Closing the Loop Between Vision and Language. Venice, Italy. 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.00812 2018-09-05 cs.CL cs.CV 81%

RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes

Semih Yagcioglu, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments EMNLP 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.10369 2018-04-18 cs.LG cs.CL cs.CV cs.IT cs.MA math.IT 81%

Emergent Communication in a Multi-Modal, Multi-Step Referential Game

Katrina Evtimova, Andrew Drozdov, Douwe Kiela, Kyunghyun Cho

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Published as a conference paper at ICLR 2018. 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.00163 2017-09-26 stat.ML cs.AI cs.CV 81%

X-CNN: Cross-modal Convolutional Neural Networks for Sparse Datasets

Petar Veličković, Duo Wang, Nicholas D. Lane, Pietro Liò

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments To appear in the 7th IEEE Symposium Series on Computational Intelligence (IEEE SSCI 2016), 8 pages, 6 figures. Minor revisions, in response to reviewers' comments

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.04350 2017-07-10 cs.CL cs.CV 81%

Imagination improves Multimodal Translation

Desmond Elliott, Ákos Kádár

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Clarified main contributions, minor correction to Equation 8, additional comparisons in Table 2, added more related work

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.02694 2017-05-09 cs.HC cs.AI cs.CV cs.LG 81%

Multimodal Affect Analysis for Product Feedback Assessment

Amol S Patwardhan, Gerald M Knapp

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, ISERC 2013, IIE Annual Conference. Proceedings. Institute of Industrial Engineers

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.08321 2016-11-28 cs.LG cs.CL cs.CV 81%

Training and Evaluating Multimodal Word Embeddings with Large-scale Web Annotated Images

Junhua Mao, Jiajing Xu, Yushi Jing, Alan Yuille

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Appears in NIPS 2016. The datasets introduced in this work will be gradually released on the project page

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17452 2026-07-21 cs.CL 新提交 80%

How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks

对话状态的多模态信号有多可靠?来自远程二元协作任务的证据

Tahiya Chowdhury

机构 * Colby College(科尔比学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 研究探讨从多模态行为测量对话状态的特征可靠性,提出三维评估框架用于视频会议二元对话特征评估,发现语言特征预测佳但跨任务通用性差,声学可靠性受说话者身份影响,交互特征是唯一可靠信号,强调相关评估对对话系统特征选择的重要性。

Comments Accepted, to appear in Proceedings of ACM International Conference on Multimodal Interaction 2026, 13 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11192 2026-07-16 cs.CV 版本更新 80%

GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

GDP.pdf:针对专业PDF文档的基础多模态推理基准测试

Suhaas Garre, Emily Ritchie, Sushant Mehta, Edwin Chen

机构 * Surge AI

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究针对专业PDF文档构建多模态推理基准测试GDP.pdf,由专业人员编写问题-文档对,通过严格筛选保留问题,有详细评分标准和能力分类。评估七个前沿模型,发现多数错误源于特定模式,公开了完整基准测试。

Comments 9 pages. v2: results updated to July 2026 leaderboard (17 models). Accepted at the 2nd Workshop on Knowledge-Intensive Multimodal Reasoning (KnowledgeMR) at CVPR 2026 (non-archival), under the former title "PDFParse: A Benchmark for Grounded Multimodal Reasoning over Professional PDF Documents". Dataset: https://huggingface.co/datasets/surgeai/GDP.pdf ; Code: https://github.com/surge-ai/gdp-pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09643 2026-04-17 cs.ET cs.AI 80%

MM-tau-p$^2$: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings

MM-tau-p$^2$: 人格自适应提示用于双控制设置中多模态代理的鲁棒性评估

Anupam Purwar, Aditya Choudhary

机构 * Sprinklr AI

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI;multimodal(comments)

AI总结 本文提出MM-tau-p$^2$基准,通过12个新指标评估多模态代理在双控制环境中的鲁棒性,结合用户输入解决查询,并展示即使使用前沿LLM,多模态鲁棒性等指标仍需考虑。

Comments A benchmark for evaluating multimodal both voice and text LLM agents in dualcontrol settings. We introduce persona adaptive prompting and 12 new metrics to assess robustness safety efficiency and recovery in customer support scenarios

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21364 2025-11-27 cs.LG cs.CV 80%

BanglaMM-Disaster: A Multimodal Transformer-Based Deep Learning Framework for Multiclass Disaster Classification in Bangla

BanglaMM-Disaster: 一种基于Transformer的多模态深度学习框架,用于孟加拉语多类灾害分类

Ariful Islam, Md Rifat Hossen, Md. Mahmudul Arif, Abdullah Al Noman, Md Arifur Rahman

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Chittagong University of Engineering and Technology(奇特格隆工程与技术大学) Department of Electronics and Telecommunication Engineering(电子与电信工程系) Wilmington University(维明顿大学) College of Graduate and Professional Studies(研究生与专业研究学院) Trine University(特林大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 BanglaMM-Disaster通过结合文本和视觉数据,提出了一种多模态深度学习框架,用于孟加拉语多类灾害分类,提升了灾害响应效率。

Comments Presented at the 2025 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON), November 21-22, 2025, University of Rajshahi, Bangladesh. 6 pages, 9 disaster classes, multimodal dataset with 5,037 samples

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02794 2025-11-05 cs.AI cs.MA 80%

When One Modality Sabotages the Others: A Diagnostic Lens on Multimodal Reasoning

Chenyu Zhang, Minsol Kim, Shohreh Ghorbani, Jingyao Wu, Rosalind Picard, Patricia Maes, Paul Pu Liang

机构 * Harvard University(哈佛大学) MIT Media Lab(麻省理工学院媒体实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at the Multimodal Algorithmic Reasoning (MAR) Workshop, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07048 2025-10-07 cs.CV 80%

Comprehensive Evaluation of Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata

Bruce Coburn, Jiangpeng He, Megan E. Rollo, Satvinder S. Dhaliwal, Deborah A. Kerr, Fengqing Zhu

机构 * Purdue University(普渡大学) Indiana University(印第安纳大学) Curtin University(Curtin大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments The extended full version of the accepted paper in 2025 IEEE BHI conference with title: Evaluating Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata. Dataset is available at: https://skynet.ecn.purdue.edu/~coburn6/ACETADA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18111 2025-05-26 cs.CV 80%

Adapting SAM 2 for Visual Object Tracking: 1st Place Solution for MMVPR Challenge Multi-Modal Tracking

Cheng-Yen Yang, Hsiang-Wei Huang, Pyong-Kun Kim, Chien-Kai Kuo, Jui-Wei Chang, Kwang-Ju Kim, Chung-I Huang, Jenq-Neng Hwang

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ICPR Multi-Modal Visual Pattern Recognition Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17492 2024-10-30 physics.chem-ph cs.AI cs.LG 80%

Unraveling Molecular Structure: A Multimodal Spectroscopic Dataset for Chemistry

Marvin Alberts, Oliver Schilter, Federico Zipoli, Nina Hartrampf, Teodoro Laino

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 29 pages, submited to conference, code available at: https://github.com/rxn4chemistry/multimodal-spectroscopic-dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15408 2024-03-26 eess.SP cs.AI cs.LG 80%

Multi-modal Heart Failure Risk Estimation based on Short ECG and Sampled Long-Term HRV

Sergio González, Abel Ko-Chun Yi, Wan-Ting Hsieh, Wei-Chao Chen, Chun-Li Wang, Victor Chien-Chia Wu, Shang-Hung Chang

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

Journal ref S. González, A. K.-C. Yi, W.-T. Hsieh, W.-C. Chen, C.-L. Wang, V. C.-C. Wu, S.-H. Chang, Multi-modal heart failure risk estimation based on short ECG and sampled long-term HRV, Information Fusion 107 (2024) 102337

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11020 2023-09-26 cs.CL cs.HC cs.RO 80%

Towards Objective Evaluation of Socially-Situated Conversational Robots: Assessing Human-Likeness through Multimodal User Behaviors

Koji Inoue, Divesh Lala, Keiko Ochi, Tatsuya Kawahara, Gabriel Skantze

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by 25th ACM International Conference on Multimodal Interaction (ICMI '23), Late-Breaking Results

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05417 2023-07-20 cs.CV 80%

The MONET dataset: Multimodal drone thermal dataset recorded in rural scenarios

Luigi Riz, Andrea Caraffa, Matteo Bortolon, Mohamed Lamine Mekhalfi, Davide Boscaini, André Moura, José Antunes, André Dias, Hugo Silva, Andreas Leonidou, Christos Constantinides, Christos Keleshis, Dante Abate, Fabio Poiesi

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Published in Computer Vision and Pattern Recognition (CVPR) Workshops 2023 - 6th Multimodal Learning and Applications Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.05542 2021-05-13 cs.CL 80%

!Qué maravilla! Multimodal Sarcasm Detection in Spanish: a Dataset and a Baseline

Khalid Alnajjar, Mika Hämäläinen

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to The Third Workshop on Multimodal Artificial Intelligence (MAI-Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.03915 2018-08-07 cs.CL cs.LG stat.ML 80%

Seq2Seq2Sentiment: Multimodal Sequence to Sequence Models for Sentiment Analysis

Hai Pham, Thomas Manzini, Paul Pu Liang, Barnabas Poczos

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments 8 pages of content, 11 pages total, 2 figures. Published as a workshop paper at ACL 2018, Proceedings of Grand Challenge and Workshop on Human Multimodal Language (Challenge-HML). 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05502 2026-07-24 cs.CV cs.AI cs.CL 版本更新 80%

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

MELLA:弥合低资源语言多模态大语言模型的语言能力与文化根基

Yufei Gao, Jiaying Fei, Nuo Chen, Ruirui Chen, Guohang Yan, Yunshi Lan, Botian Shi

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) East China Normal University(东华大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Institute of High Performance Computing, A*STAR(高性能计算研究所,A*STAR)

专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究针对低资源语言MLLMs文化内涵不足问题,提出MELLA数据集,采用双源策略,结合母语网络图像-替代文本对与生成翻译的图像描述进行监督,经实验表明能减轻文化幻觉,强调数据对齐对低资源语言文化基础多模态理解的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04413 2026-02-05 cs.CL cs.AI cs.MM 80%

History-Guided Iterative Visual Reasoning with Self-Correction

基于历史的迭代视觉推理与自我校正

Xinglong Yang, Zhilin Peng, Zhanzhan Liu, Haochen Shi, Sheng-Jun Huang

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 H-GIVR框架通过迭代视觉推理与自我校正,显著提升多模态推理准确性并保持低计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17856 2025-05-02 cs.LG eess.SP 80%

Enhancing clinical decision support with physiological waveforms -- a multimodal benchmark in emergency care

Juan Miguel Lopez Alcaraz, Hjalmar Bouma, Nils Strodthoff

专题命中 多模态评测 :multimodal(title,abstract)

Comments Version accepted by Computers in Biology and Medicine: 21 pages, 2 figures, code available under https://github.com/AI4HealthUOL/MDS-ED, dataset available under https://physionet.org/content/multimodal-emergency-benchmark/

Journal ref J.M. Lopez Alcaraz, H. Bouma, N. Strodthoff, Enhancing clinical decision support with physiological waveforms -- A multimodal benchmark in emergency care, Computers in Biology and Medicine, Vol. 192, Part A, 2025, 110196

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12219 2024-10-17 cs.AI cs.CL cs.MM 80%

OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities

Lichang Chen, Hexiang Hu, Mingda Zhang, Yiwen Chen, Zifeng Wang, Yandong Li, Pranav Shyam, Tianyi Zhou, Heng Huang, Ming-Hsuan Yang, Boqing Gong

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);omni-modal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments 19 pages, 6 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.10504 2021-09-06 eess.IV 80%

Liver Segmentation from Multimodal Images using HED-Mask R-CNN

Supriti Mulay, Deepika G, Jeevakala S, Keerthi Ram, Mohanasankar Sivaprakasam

专题命中 多模态评测 :multimodal(title,abstract)

Comments Accepted in 1st International Workshop on Multiscale Multimodal Medical Imaging (MMMI 2019) - MICCAI 2019

Journal ref Multiscale Multimodal Medical Imaging. MMMI 2019. Lecture Notes in Computer Science, vol 11977

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17884 2026-08-19 cs.CV 新提交 79%

CFB-GBM v2.0: An Augmented Longitudinal Dataset for Multi-Modal Glioblastoma Segmentation, Radiomics, and RANO Progression Tracking

CFB-GBM v2.0:用于多模态胶质母细胞瘤分割、放射组学及RANO进展追踪的增强纵向数据集

Alexandre G. Leclercq, Noémie N. Moreau, Hugo Audebert, Andros Nassar, Thomas Cochin, Thomas Leleu, Loïc Le Henaff, Alexis Desmonts, Yoann Poirier, Aurélie Dubru, Laura Guillemette, Pascal Lecoeur, Kévin Lemasson, Cyril Jaudet, Sébastien Bougleux, Romain Hérault, Carole Brunaud, Samuel Valable, Dinu Stefan, Charlotte Raboutet, Alain Batalla, Joëlle Lacroix, Roman Rouzier, Aurélien Corroyer-Dulmont

机构 * Centre François Baclesse(弗朗索瓦·巴克莱斯中心) Université de Caen Normandie(卡昂诺曼底大学) ENSICAEN(卡昂高等工程师学院) GREYC(格雷计算机科学研究中心)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文发布增强版CFB-GBM v2.0纵向数据集,含264名GBM患者数据,完成所有时间点GTV勾画(完成率达97%),提供相关标注、特征及WHO分类信息,可用于多模态GBM相关研究。

Comments 9 pages, 2 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08093 2026-08-19 cs.AI 版本更新 79%

A Multimodal Agentic Pathology Co-pilot via Evidence Grounded Reasoning

面向证据基础计算病理学的多模态智能体协同助手

Zhe Xu, Zhengyu Zhang, Zhiyuan Cai, Jiahao Xu, Yijie Lin, Ziyi Liu, Junlin Hou, Hongyi Wang, Yuxiang Nie, Yihui Wang, Jiabo Ma, Ling Liang, Yingxue Xu, Zhengrui Guo, Guanghao Wu, Danyi Li, Ziqi Zhou, Donglin Tan, Zhijian Cen, Ying Tan, Xiaolin Liu, Qi Xie, Xiaoying Tang, Xi Peng, Cheng Deng, Lijuan Qu, Ronald Cheong Kin Chan, Li Liang, Hao Chen

机构 * Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Department of Pathology, Nanfang Hospital, Southern Medical University(南方医科大学南芳医院病理科) Department of Pathology, School of Basic Medical Sciences, Southern Medical University(南方医科大学基础医学学院病理科) Department of Anatomical and Cellular Pathology, Chinese University of Hong Kong(香港中文大学解剖与细胞病理学系) Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理学重点实验室) Jinfeng Laboratory(锦风实验室) Department of Chemical and Biological Engineering, Hong Kong University of Science and Technology(香港科技大学化学与生物工程系) Division of Life Science, Hong Kong University of Science and Technology(香港科技大学生命科学系) State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology(香港科技大学深圳-香港协同创新研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 提出PathPocket,一种多模态AI协同助手,通过构建包含11万文档的病理证据语料库和455万实体的超图,实现基于证据的病理诊断,在20万真实案例上超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏