arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4868 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4868 篇

2506.18201 2025-09-05 cs.CL cs.CV cs.HC 81%

Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications

Bushra Asseri, Estabraq Abdelaziz, Maha Al Mogren, Tayef Alhefdhi, Areej Al-Wabil

机构 * College of Engineering & Advanced Computing(工程与高级计算学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17860 2025-08-26 cs.CV cs.AI 81%

AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering

Kang Zeng, Guojin Zhong, Jintao Cheng, Jin Yuan, Zhiyong Li

机构 * Hunan University(湖南大学) South China Normal University(华南师范大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09195 2025-08-14 eess.IV cs.AI cs.CV 81%

impuTMAE: Multi-modal Transformer with Masked Pre-training for Missing Modalities Imputation in Cancer Survival Prediction

Maria Boyko, Aleksandra Beliaeva, Dmitriy Kornilov, Alexander Bernstein, Maxim Sharaev

机构 * Center for Applied AI, Skolkovo Institute of Science and Technology, Moscow, Russian Federation(应用人工智能中心,斯克尔科夫科学与技术研究所,莫斯科,俄罗斯联邦) BIMAI-Lab, Biomedically Informed Artificial Intelligence Laboratory, University of Sharjah, Sharjah, United Arab Emirates(BIMAI实验室,生物医学导向人工智能实验室,肖贾联合大学,肖贾,阿拉伯联合酋长国) Ivannikov Institute for System Programming of the Russian Academy of Sciences, Moscow, Russian Federation(伊万诺夫俄罗斯科学院系统编程研究所,莫斯科,俄罗斯联邦)

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09175 2025-08-14 cs.CV cs.AI 81%

A Context-aware Attention and Graph Neural Network-based Multimodal Framework for Misogyny Detection

Mohammad Zia Ur Rehman, Sufyaan Zahoor, Areeb Manzoor, Musharaf Maqbool, Nagendra Kumar

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Indian Institute of Technology Indore(印度理工学院印德分校) Department of Computer Science and Engineering, National Institute of Technology Srinagar(计算机科学与工程系,印度国立技术学院斯里 Nagar分校)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published in Information Processing & Management

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20670 2025-07-29 cs.CV cs.AI cs.LG 81%

A Multimodal Architecture for Endpoint Position Prediction in Team-based Multiplayer Games

Jonas Peche, Aliaksei Tsishurou, Alexander Zap, Guenter Wallner

机构 * Computer Graphics(计算机图形学) Johannes Kepler University Linz(约瑟夫·施特劳斯大学林茨) DS Research(DS研究) Wargaming(沃拉基姆)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07783 2025-07-28 cs.CV cs.CL 81%

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Zhaokai Wang, Xizhou Zhu, Xue Yang, Gen Luo, Hao Li, Changyao Tian, Wenhan Dou, Junqi Ge, Lewei Lu, Yu Qiao, Jifeng Dai

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Sensetime

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Journal ref TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09647 2025-07-18 cs.MM cs.AI 81%

KEN: Knowledge Augmentation and Emotion Guidance Network for Multimodal Fake News Detection

Peican Zhu, Yubo Jing, Le Cheng, Keke Tang, Yangming Guo

机构 * Northwestern Polytechnical University(西北工业大学) Guangzhou University(广东大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23121 2025-07-15 eess.IV cs.AI cs.CV cs.LG 81%

CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation

Xinlei Yu, Changmiao Wang, Hui Jin, Ahmed Elazab, Gangyong Jia, Xiang Wan, Changqing Zou, Ruiquan Ge

机构 * Hangzhou Dianzi University(杭州电子科技大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) Shenzhen University(深圳大学) Zhejiang University(浙江大学)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted By ACMMM25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06744 2025-07-10 cs.CV cs.LG cs.MM 81%

Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching

Yafei Zhang, Yongle Shang, Huafeng Li

机构 * Kunming University of Science and Technology(昆明理工大学)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00185 2025-07-02 eess.IV cs.AI cs.CV 81%

Multimodal, Multi-Disease Medical Imaging Foundation Model (MerMED-FM)

Yang Zhou, Chrystie Wan Ning Quek, Jun Zhou, Yan Wang, Yang Bai, Yuhe Ke, Jie Yao, Laura Gutierrez, Zhen Ling Teo, Darren Shu Jeng Ting, Brian T. Soetikno, Christopher S. Nielsen, Tobias Elze, Zengxiang Li, Linh Le Dinh, Lionel Tim-Ee Cheng, Tran Nguyen Tuan Anh, Chee Leong Cheng, Tien Yin Wong, Nan Liu, Iain Beehuat Tan, Tony Kiat Hon Lim, Rick Siow Mong Goh, Yong Liu, Daniel Shu Wei Ting

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 42 pages, 3 composite figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21600 2025-06-30 cs.CL cs.AI cs.IR 81%

Structured Attention Matters to Multimodal LLMs in Document Understanding

Chang Liu, Hongkai Chen, Yujun Cai, Hang Wu, Qingwen Ye, Ming-Hsuan Yang, Yiwei Wang

机构 * vivo Mobile Communication Co., Ltd(vivo移动通信有限公司) The University of Queensland(昆士兰大学) University of California, Merced(加州大学默塞德分校)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17364 2025-06-25 cs.CY cs.AI cs.CV cs.HC 81%

AI-based Multimodal Biometrics for Detecting Smartphone Distractions: Application to Online Learning

Alvaro Becerra, Roberto Daza, Ruth Cobos, Aythami Morales, Mutlu Cukurova, Julian Fierrez

机构 * Universidad Autónoma de Madrid(马德里自治大学) University College London(伦敦大学学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in EC-TEL25: 20th European Conference on Technology Enhanced Learning, Newcastle and Durham, UK, 15-19 September 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13360 2025-06-04 cs.CV cs.AI cs.LG 81%

Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning

Hai-Long Sun, Zhun Sun, Houwen Peng, Han-Jia Ye

机构 * School of Artificial Intelligence, Nanjing University(南京大学人工智能学院) National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Tencent(腾讯) Center for Language AI Research, Tohoku University(东北大学语言人工智能研究中心) RIKEN Center for Advanced Intelligence Project(日本理化学研究所先进人工智能项目中心)

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ACL 2025. The project page is available at https://sun-hailong.github.io/projects/TVC

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16192 2025-06-02 cs.CV cs.AI 81%

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Chaoya Jiang, Yongrui Heng, Wei Ye, Han Yang, Haiyang Xu, Ming Yan, Ji Zhang, Fei Huang, Shikun Zhang

机构 * National Engineering Research Center for Software Engineering, Peking University(软件工程国家工程研究中心,北京大学) Alibaba Group(阿里巴巴集团) ZEEKR Intelligent Technology Holding Limited(ZEKR智能科技控股有限公司)

专题命中 其他多模态 :multimodal(title);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24026 2025-06-02 cs.CV cs.AI 81%

MaskAdapt: Unsupervised Geometry-Aware Domain Adaptation Using Multimodal Contextual Learning and RGB-Depth Masking

Numair Nadeem, Muhammad Hamza Asad, Saeed Anwar, Abdul Bais

机构 * University of Regina(里贾纳大学) University Canada West(加拿大西部大学) The University of Western Australia(西澳大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 11 pages, 5 figures, presented at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025. Reviewer comments available upon request

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03122 2025-05-22 cs.CL cs.AI 81%

The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models

Zichao Li, Xueru Wen, Jie Lou, Yuqiu Ji, Yaojie Lu, Xianpei Han, Debing Zhang, Le Sun

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17787 2025-05-20 cs.AI cs.CR cs.CV 81%

Evaluating the Efficacy of Prompt-Engineered Large Multimodal Models Versus Fine-Tuned Vision Transformers in Image-Based Security Applications

Fouad Trad, Ali Chehab

机构 * Electrical and Computer Engineering(电气与计算机工程)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref Published in ACM Transactions on Intelligent Systems and Technology, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20291 2025-05-16 cs.CV cs.AI cs.LG q-bio.BM 81%

CryoSAMU: Enhancing 3D Cryo-EM Density Maps of Protein Structures at Intermediate Resolution with Structure-Aware Multimodal U-Nets

Chenwei Zhang, Khanh Dao Duc

机构 * Department of Computer Science, UBC(计算机科学系,不列颠哥伦比亚大学) Department of Mathematics, UBC(数学系,不列颠哥伦比亚大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 19 pages, 6 main figures, 2 supplementary figures, 3 main tables, 4 supplementary tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13884 2025-04-22 cs.HC cs.AI cs.CV 81%

Towards a Multimodal Document-grounded Conversational AI System for Education

Karan Taneja, Anjali Singh, Ashok K. Goel

机构 * Georgia Institute of Technology(佐治亚理工学院) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 15 pages, 4 figures, AIED 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21834 2025-03-31 cs.CV cs.AI 81%

A Multi-Modal Knowledge-Enhanced Framework for Vessel Trajectory Prediction

Haomin Yu, Tianyi Li, Kristian Torp, Christian S. Jensen

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18403 2025-03-25 cs.CV cs.AI cs.LG 81%

Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning

Xusheng Cao, Haori Lu, Linlan Huang, Fei Yang, Xialei Liu, Ming-Ming Cheng

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11911 2025-03-25 cs.LG cs.AI cs.CV cs.RO 81%

ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling

Zikang Zhou, Hengjian Zhou, Haibo Hu, Zihao Wen, Jianping Wang, Yung-Hui Li, Yu-Kai Huang

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12326 2025-03-18 cs.CV cond-mat.mtrl-sci cs.AI 81%

Leveraging Vision Capabilities of Multimodal LLMs for Automated Data Extraction from Plots

Maciej P. Polak, Dane Morgan

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00915 2025-03-04 cs.CV cs.AI 81%

Multimodal Distillation-Driven Ensemble Learning for Long-Tailed Histopathology Whole Slide Images Analysis

Xitong Ling, Yifeng Ping, Jiawen Li, Jing Peng, Yuxuan Chen, Minxi Ouyang, Yizhi Wang, Yonghong He, Tian Guan, Xiaoping Liu, Lianghui Zhu

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13383 2025-02-20 cs.CL cs.CV cs.LG 81%

MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification

Linzhuang Sun, Hao Liang, Jingxuan Wei, Bihui Yu, Tianpeng Li, Fan Yang, Zenan Zhou, Wentao Zhang

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11573 2025-02-18 cs.CL cs.AI 81%

InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning

Congkai Xie, Shuo Cai, Wenjun Wang, Pengxiang Li, Zhijie Sang, Kejing Yang, Yiming Zhang, Zhen Li, Guanghao Zhu, Zeyu Liu, Yang Yu, Yuhang Liu, Su Lu, Baoyi He, Qi Zhou, Xiaotian Han, Jianbo Yuan, Shengyu Zhang, Fei Wu, Hongxia Yang

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10810 2025-02-18 cs.CV cs.AI cs.AR cs.LG 81%

FabGPT: An Efficient Large Multimodal Model for Complex Wafer Defect Knowledge Queries

Yuqi Jiang, Xudong Lu, Qian Jin, Qi Sun, Hanming Wu, Cheng Zhuo

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published in ACM/IEEE International Conference On Computer Aided Design (ICCAD) 2024. Corresponding Author: Qi Sun (qisunchn@zju.edu.cn)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13607 2025-01-28 cs.CV cs.AI 81%

MM-NeRF: Multimodal-Guided 3D Multi-Style Transfer of Neural Radiance Field

Zijiang Yang, Zhongwei Qiu, Chang Xu, Dongmei Fu

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published in: IEEE Transactions on Visualization and Computer Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10377 2025-01-17 eess.IV cs.AI cs.CV 81%

Enhanced Masked Image Modeling to Avoid Model Collapse on Multi-modal MRI Datasets

Linxuan Han, Sa Xiao, Zimeng Li, Haidong Li, Xiuchao Zhao, Yeqing Han, Fumin Guo, Xin Zhou

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments This work has been submitted to the lEEE for possible publication. copyright may be transferred without notice, after which this version may no longer be accessible

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09784 2024-12-16 cs.CL cs.AI 81%

Semi-IIN: Semi-supervised Intra-inter modal Interaction Learning Network for Multimodal Sentiment Analysis

Jinhao Lin, Yifei Wang, Yanwu Xu, Qi Liu

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏