arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9111 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9111 篇

2404.09992 2025-07-01 cs.CV cs.AI cs.CL 82%

MMInA: Benchmarking Multihop Multimodal Internet Agents

Shulin Tian, Ziniu Zhang, Liangyu Chen, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2025 findings. The live leaderboard is at https://mmina.cliangyu.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21523 2025-06-23 cs.CL cs.AI cs.CV 82%

More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

Chengzhi Liu, Zhongxing Xu, Qingyue Wei, Juncheng Wu, James Zou, Xin Eric Wang, Yuyin Zhou, Sheng Liu

机构 * UC Santa Cruz(加州大学圣克ruz分校) Stanford University(斯坦福大学) UC Santa Barbara(加州大学圣芭芭拉分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10855 2025-06-23 cs.CL cs.AI cs.CV 82%

Core Knowledge Deficits in Multi-Modal Language Models

Yijiang Li, Qingying Gao, Tianwei Zhao, Bingyang Wang, Haoran Sun, Haiyun Lyu, Robert D. Hawkins, Nuno Vasconcelos, Tal Golan, Dezhi Luo, Hokin Deng

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ICML 2025. Project page at https://williamium3000.github.io/core-knowledge and code is available at https://github.com/williamium3000/core-knowledge

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06752 2025-06-18 cs.RO 82%

Semantic Enhancement for Object SLAM with Heterogeneous Multimodal Large Language Model Agents

Jungseok Hong, Ran Choi, John J. Leonard

机构 * Computer Science and Artificial Intelligence Laboratory (CSAIL) at the Massachusetts Institute of Technology (MIT)(麻省理工学院计算机科学与人工智能实验室)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10376 2025-06-13 cs.SE cs.HC 82%

MLLM-Based UI2Code Automation Guided by UI Layout Information

Fan Wu, Cuiyun Gao, Shuqing Li, Xin-Cheng Wen, Qing Liao

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract)

Comments Accepted by the 34th International Symposium on Software Testing and Analysis (ISSTA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10359 2025-06-13 cs.RO cs.LG 82%

Demonstrating Multi-Suction Item Picking at Scale via Multi-Modal Learning of Pick Success

Che Wang, Jeroen van Baar, Chaitanya Mitash, Shuai Li, Dylan Randle, Weiyao Wang, Sumedh Sontakke, Kostas E. Bekris, Kapil Katyal

机构 * Amazon Robotics(亚马逊机器人技术)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

Comments Accepted to Robotics: Science and Systems (RSS 2025), 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11300 2025-06-10 cs.CL cs.AI cs.CV 82%

CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?

Aashish Anantha Ramakrishnan, Aadarsh Anantha Ramakrishnan, Dongwon Lee

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) National Institute of Technology, Tiruchirappalli(特里奇里帕利学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments To appear at the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), Vienna, Austria, July 2025, https://2025.aclweb.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.20331 2025-06-10 cs.CV cs.AI cs.CL cs.LG 82%

Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models

Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang, Yifei Ming, Qing Yu, Go Irie, Yixuan Li, Hai Li, Ziwei Liu, Kiyoharu Aizawa

机构 * The University of Tokyo(东京大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室) Duke University(杜克大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) LY Corporation(LY公司) Tokyo University of Science(东京科学大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ACL 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05523 2025-06-09 cs.CV cs.AI cs.CL cs.LG 82%

MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning

Zikui Cai, Andrew Wang, Anirudh Satheesh, Ankit Nakhawa, Hyunwoo Jae, Keenan Powell, Minghui Liu, Neel Jay, Sungbin Oh, Xiyao Wang, Yongyuan Liang, Tom Goldstein, Furong Huang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04688 2025-06-06 cs.CL cs.AI cs.CV 82%

MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models

Gio Paik, Geewook Kim, Jinbae Im

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00958 2025-06-03 cs.AI cs.CL cs.CV 82%

Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues

Youngmin Kim, Jiwan Chung, Jisoo Kim, Sunghyun Lee, Sangkyu Lee, Junhyeok Kim, Cheoljong Yang, Youngjae Yu

机构 * Yonsei University(延世大学) NC Research, NCSOFT Corporation(NC研究,NCSOFT公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ACL 2025 (Main), Our code and dataset: https://github.com/winston1214/nonverbal-conversation

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00421 2025-06-03 cs.CL cs.AI cs.CV 82%

Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions

Jihyoung Jang, Minwook Bae, Minji Kim, Dilek Hakkani-Tur, Hyounghun Kim

机构 * Graduate School of Artificial Intelligence, POSTECH(POSTECH人工智能研究生院) Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系) Artificial Intelligence Graduate School, UNIST(UNIST人工智能研究生院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2025 (32 pages); Project website: https://m3c-dataset.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11515 2025-06-03 cs.CV cs.CL cs.MM 82%

Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment

Kangsheng Wang, Chengwei Ye, Huanzhen Zhang, Linuo Xu, Shuyan Liu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments The article contains serious scientific errors and cannot be corrected by updating the preprint

Journal ref IECE Trans. Emerg. Top. Artif. Intell. 2 (2025) 57--67

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18117 2025-06-02 q-bio.NC q-bio.QM 82%

Multi-Modal Spectral Parametrization Method (MMSPM) for analyzing EEG activity with distinct scaling regimes

Frigyes Samuel Racz, John Milton, Juan Luis Cabrera, Gábor Csukly, José del R. Millán

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

Comments 6 pages, 4, figures. This work has been submitted for publication to the 12th Annual IEEE EMBS International Conference on Neural Engineering (https://neuro.embs.org/2025/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01081 2025-05-22 cs.CV cs.AI cs.CL 82%

The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles

Vernon Y. H. Toh, Yew Ken Chia, Deepanway Ghosal, Soujanya Poria

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13483 2025-05-21 cs.CL cs.AI cs.CV 82%

EmoMeta: A Multimodal Dataset for Fine-grained Emotion Classification in Chinese Metaphors

Xingyuan Lu, Yuxi Liu, Dongyu Zhang, Zhiyao Wu, Jing Ren, Feng Xia

机构 * School of Software(软件学院) Dalian University of Technology(大连理工大学) School of Foreign Languages and School of Software(外语学院和软件学院) Faculty of Business Administration(商学院) University of Macau(澳门大学) School of Computing Technologies(计算技术学院) RMIT University(皇家墨尔本理工大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07266 2025-05-13 cs.RO 82%

BETTY Dataset: A Multi-modal Dataset for Full-Stack Autonomy

Micah Nye, Ayoub Raji, Andrew Saba, Eidan Erlich, Robert Exley, Aragya Goyal, Alexander Matros, Ritesh Misra, Matthew Sivaprakasam, Marko Bertogna, Deva Ramanan, Sebastian Scherer

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Waterloo(滑铁卢大学) University of Pittsburgh(匹兹堡大学)

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract)

Comments 8 pages. 5 figures. ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04653 2025-05-09 cs.CL cs.AI cs.CV cs.LG 82%

Advancing Conversational Diagnostic AI with Multimodal Reasoning

Khaled Saab, Jan Freyberg, Chunjong Park, Tim Strother, Yong Cheng, Wei-Hung Weng, David G. T. Barrett, David Stutz, Nenad Tomasev, Anil Palepu, Valentin Liévin, Yash Sharma, Roma Ruparel, Abdullah Ahmed, Elahe Vedadi, Kimberly Kanada, Cian Hughes, Yun Liu, Geoff Brown, Yang Gao, Sean Li, S. Sara Mahdavi, James Manyika, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Pushmeet Kohli, S. M. Ali Eslami, Joëlle Barral, Adam Rodman, Vivek Natarajan, Mike Schaekermann, Tao Tu, Alan Karthikesalingam, Ryutaro Tanno

机构 * Google(谷歌) DeepMind(深Mind) Google Research(谷歌研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16566 2025-05-08 cs.HC 82%

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

Zheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun, Licai Sun, Yong Ren, Zebang Cheng, Bin Liu, Rui Liu, Xiaojiang Peng, Jiangyan Yi, Jianhua Tao

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01794 2025-05-06 cs.CL cs.AI cs.MM 82%

A Multimodal Framework for Explainable Evaluation of Soft Skills in Educational Environments

Jared D. T. Guerrero-Sosa, Francisco P. Romero, Víctor Hugo Menéndez-Domínguez, Jesus Serrano-Guerrero, Andres Montoro-Montarroso, Jose A. Olivas

机构 * University of Castilla-La Mancha(卡斯蒂利亚-拉曼查大学) Autonomous University of Yucatan(尤卡坦自治大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16427 2025-04-25 cs.CL cs.AI cs.MM 82%

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

Hanlei Zhang, Zhuohang Li, Yeshuang Zhu, Hua Xu, Peiwu Wang, Haige Zhu, Jie Zhou, Jinchao Zhang

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯人工智能研究院) Kennesaw State University(凯斯西储大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments 23 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09354 2025-04-15 cs.CV cs.AI cs.CL cs.LG q-bio.QM 82%

REMEMBER: Retrieval-based Explainable Multimodal Evidence-guided Modeling for Brain Evaluation and Reasoning in Zero- and Few-shot Neurodegenerative Diagnosis

Duy-Cat Can, Quang-Huy Tang, Huong Ha, Binh T. Nguyen, Oliver Y. Chén

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02217 2025-04-04 cs.HC 82%

The Plot Thickens: Quantitative Part-by-Part Exploration of MLLM Visualization Literacy

Matheus Valentim, Vaishali Dhanoa, Gabriela Molina León, Niklas Elmqvist

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract)

Comments 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01767 2025-04-03 eess.AS cs.AI cs.CV 82%

Leveraging Embedding Techniques in Multimodal Machine Learning for Mental Illness Assessment

Abdelrahaman A. Hassan, Abdelrahman A. Ali, Aya E. Fouda, Radwa J. Hanafy, Mohammed E. Fouda

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16081 2025-03-31 cs.LG cs.IR 82%

OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning

Zhiyuan Liu, Yuting Zhang, Feng Liu, Changwang Zhang, Ying Sun, Jun Wang

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08182 2025-03-21 cs.CV cs.AI cs.CL 82%

MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models

Wenbo Hu, Jia-Chen Gu, Zi-Yi Dou, Mohsen Fayyaz, Pan Lu, Kai-Wei Chang, Nanyun Peng

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17250 2025-03-20 cs.CL cs.AI cs.CV 82%

JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation

Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira, Jeonghun Baek, Xiang Yue, Graham Neubig, Kiyoharu Aizawa

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at NAACL 2025. Project page: https://mmmu-japanese-benchmark.github.io/JMMMU/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16643 2025-03-19 cs.CL cs.AI cs.SD eess.AS 82%

An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue

Koji Inoue, Divesh Lala, Mikey Elmers, Keiko Ochi, Tatsuya Kawahara

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments This paper has been accepted for presentation at International Workshop on Spoken Dialogue Systems Technology 2025 (IWSDS 2025) and represents the author's version of the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13399 2025-03-18 cs.CV cs.AI cs.CL cs.LG q-bio.CB 82%

MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research

James Burgess, Jeffrey J Nirschl, Laura Bravo-Sánchez, Alejandro Lozano, Sanket Rajan Gupte, Jesus G. Galaz-Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, Sarina M. Hasan, Alexandra Johannesson, William D. Leineweber, Malvika G Nair, Ridhi Yarlagadda, Connor Zuraski, Wah Chiu, Sarah Cohen, Jan N. Hansen, Manuel D Leonetti, Chad Liu, Emma Lundberg, Serena Yeung-Levy

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR 2025 (Conference on Computer Vision and Pattern Recognition) Project page at https://jmhb0.github.io/microvqa Benchmark at https://huggingface.co/datasets/jmhb/microvqa

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10627 2025-03-14 cs.CV cs.AI cs.CL 82%

SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems

Ziyu Guo, Ray Zhang, Hao Chen, Jialin Gao, Dongzhi Jiang, Jiaze Wang, Pheng-Ann Heng

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Initially released in September 2024. Project page: https://sciverse-cuhk.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏