arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4633 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4633 篇

2601.09298 2026-05-08 cs.CV 89%

Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain

基于多模态LLM的ICT图像描述:弥合通用与行业领域之间的差距

Lianying Chao, Kai Zhang, Haoran Cai, Sijie Wu, Xubin Li, Xin Chen

机构 * GTS, Huawei Technologies Co., Ltd.(华为技术有限公司GTS部门)

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract,abstract_cn);multimodal(abstract);image-text(abstract)

AI总结 本文提出多阶段训练策略,构建ICT领域图像描述模型DICModel,通过合成图像文本对提升模型性能,实验表明其在BLEU指标和准确率上均优于现有模型。

Journal ref 2025 CCF BigData

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24904 2026-07-29 cs.CV cs.CL 新提交 88%

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Mage-VL:一种高效的编解码器原生流式多模态基础模型

Senqiao Yang, Kaichen Zhang, Zhaoyang Jia, Jinghao Guo, Yifei Shen, Xinjie Zhang, Xiaoyi Zhang, Haoqing Wang, Xiao Li, Peng Zhang, Xiang An, Yin Xie, Zhening Liu, Xun Guo, Jiahao Li, Shicheng Zheng, Jinglu Wang, Zongyu Guo, Wenxuan Xie, Zihan Zheng, Yuxuan Luo, Bin Li, Yan Lu

机构 * Microsoft(微软)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title);image-text(abstract);分类 cs.CV、cs.CL

AI总结 研究针对标准视觉语言模型在流式感知任务的不足,提出Mage-VL。核心方法是用Mage-ViT选择性编码关键区域,减少视觉令牌消耗。该模型在多模态理解和交互上表现出色,推理速度加快,还有多项关键实证发现。

Comments Project page: https://microsoft.github.io/Mage

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26501 2026-05-27 cs.CV cs.AI 88%

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

揭示视觉-语言模型的脆弱性:通过纹理约束扰动和跨模态优化的多模态对抗协同

Xiang Fang, Wanlong Fang, Changshuo Wang

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Nanyang Technological University, Singapore(新加坡南洋理工大学) University College London(伦敦大学学院)

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 提出多模态对抗协同框架,通过纹理约束的通用对抗扰动和可学习的文本提示扰动,在黑盒设置下联合优化,揭示视觉-语言模型在多模态攻击下的脆弱性。

Comments Publish in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14889 2025-02-24 cs.CV cs.AI 88%

Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations Interpretability

Zhiyu Zhu, Zhibo Jin, Jiayu Zhang, Nan Yang, Jiahao Huang, Jianlong Zhou, Fang Chen

专题命中 图文多模态 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09645 2025-02-17 cs.CL cs.AI 88%

From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models

Mayank Vatsa, Aparna Bharati, Surbhi Mittal, Richa Singh

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18185 2025-01-28 cs.CV cs.AI 88%

TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization

Yucong Luo, Mingyue Cheng, Jie Ouyang, Xiaoyu Tao, Qi Liu

专题命中 图文多模态 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CV、cs.AI

Comments Need a lot of refinements

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09273 2024-11-15 cs.CL cs.AI 88%

Cross-Modal Consistency in Multimodal Large Language Models

Xiang Zhang, Senyu Li, Ning Shi, Bradley Hauer, Zijun Wu, Grzegorz Kondrak, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12821 2024-08-26 cs.CV cs.AI 88%

Examining the Commitments and Difficulties Inherent in Multimodal Foundation Models for Street View Imagery

Zhenyuan Yang, Xuhui Lin, Qinyi He, Ziye Huang, Zhengliang Liu, Hanqi Jiang, Peng Shu, Zihao Wu, Yiwei Li, Stephen Law, Gengchen Mai, Tianming Liu, Tao Yang

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08515 2024-07-15 cs.CV cs.AI 88%

15M Multimodal Facial Image-Text Dataset

Dawei Dai, YuTang Li, YingGe Liu, Mingming Jia, Zhang YuanHui, Guoyin Wang

专题命中 图文多模态 :image-text(title,abstract);multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10208 2024-04-03 cs.CV cs.CL 88%

MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Changyao Tian, Xizhou Zhu, Yuwen Xiong, Weiyun Wang, Zhe Chen, Wenhai Wang, Yuntao Chen, Lewei Lu, Tong Lu, Jie Zhou, Hongsheng Li, Yu Qiao, Jifeng Dai

专题命中 图文多模态 :multi-modal(title,abstract);image-text(title,abstract);分类 cs.CV、cs.CL

Comments 20 pages, 9 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17049 2024-04-02 cs.CV cs.CL cs.LG 88%

MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training

Pavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli, Oncel Tuzel

专题命中 图文多模态 :multi-modal(title,abstract);image-text(title,abstract);分类 cs.CV、cs.CL

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05261 2024-03-29 cs.CV cs.MM 88%

Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval

Hailang Huang, Zhijie Nie, Ziqiao Wang, Ziyu Shang

专题命中 图文多模态 :cross-modal(title,abstract);image-text(title,abstract);分类 cs.CV、cs.MM

Comments 9 pages, Accepted by AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02677 2024-03-06 cs.CV cs.CL 88%

Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters

Weizhi Wang, Khalil Mrini, Linjie Yang, Sateesh Kumar, Yu Tian, Xifeng Yan, Heng Wang

专题命中 图文多模态 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CV、cs.CL

Comments Project Website: https://mlm-filter.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16575 2024-01-31 cs.CL cs.CV 88%

Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking

Ivana Beňová, Jana Košecká, Michal Gregor, Martin Tamajka, Marcel Veselý, Marián Šimko

专题命中 图文多模态 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CV、cs.CL

Comments 9 pages of text, 11 pages total, 7 figures, 3 tables, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.08354 2016-12-28 cs.CV cs.CL cs.LG 88%

Image-Text Multi-Modal Representation Learning by Adversarial Backpropagation

Gwangbeen Park, Woobin Im

专题命中 图文多模态 :multi-modal(title,abstract);image-text(title,abstract);分类 cs.CV、cs.CL

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03953 2026-04-07 cs.CV cs.LG 88%

Multimodal Structure Learning: Disentangling Shared and Specific Topology via Cross-Modal Graphical Lasso

多模态结构学习:通过跨模态图拉索解耦共享和特定拓扑

Fei Wang, Yutong Zhang, Xiong Wang

机构 * Stony Brook University(石溪大学) Sichuan University(四川大学) USTC(中国科学技术大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出跨模态图拉索方法,通过统一视觉语言编码器和跨注意力蒸馏机制,解耦共享与类别特定拓扑,提升多模态表征的可解释性。

Comments Submitted to a conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23067 2026-03-26 cs.CV 88%

MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding

MLLM-HWSI: 一种用于分层全滑动图像理解的多模态大语言模型

Basit Alawode, Arif Mahmood, Muaz Khalifa Al-Radi, Shahad Albastaki, Asim Khan, Muhammad Bilal, Moshira Ali Abdalla, Mohammed Bennamoun, Sajid Javed

机构 * Department of Computer Science, Khalifa University of Science and Technology(卡利法科技大学计算机科学系) Information Technology University(信息技术大学) KAU(卡乌大学) University of the Western Australia(西澳大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV

AI总结 本文提出MLLM-HWSI,一种分层全滑动图像级多模态大语言模型,通过四级尺度对齐视觉特征与病理语言,提升解释性证据接地推理能力,在六个CPath任务上取得新SOTA结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11680 2025-12-15 cs.CV 88%

Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing

跨模态上下文感知学习:用于遥感中视觉提示引导的多模态图像理解

Xu Zhang, Jiabin Fang, Zhuoming Ding, Jin Yuan, Xuan Liu, Qianjun Zhang, Zhiyong Li

机构 * College of Computer Science and Electronic Engineering, Hunan University(计算机科学与电子工程学院,湖南大学) School of Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University(机器人学院及机器人视觉感知与控制技术国家工程研究中心,湖南大学) School of Computing and Artificial Intelligence, Southwest Jiaotong University(计算与人工智能学院,西南交通大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

AI总结 CLV-Net通过跨模态上下文感知学习,利用视觉提示引导遥感多模态图像理解,提升目标识别精度和用户意图对齐能力。

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18751 2025-12-04 cs.CL 88%

Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and Fusion

基于分布的特征恢复与融合的鲁棒多模态图像-文本对情感分析

Daiqing Wu, Dongbao Yang, Yu Zhou, Can Ma

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) TMCC, College of Computer Science, Nankai University(TMCC,南开大学计算机学院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CL

AI总结 本文提出DRF方法,通过特征队列和分布估计,实现对图像-文本对中低质量和缺失模态的鲁棒情感分析。

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23184 2025-10-28 cs.CV 88%

Finding 3D Scene Analogies with Multimodal Foundation Models

Junho Kim, Young Min Kim

机构 * Institute of New Media and Communications(新媒体与通讯研究所) Dept. of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

Comments Accepted to FM4RoboPlan workshop at RSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04351 2025-10-14 cs.RO cs.AI 88%

MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection

Liman Wang, Hanyang Zhong, Tianyuan Wang, Shan Luo, Jihong Zhu

机构 * School of Physics, Engineering and Technology, University of York(物理、工程与技术学院,约克大学) Department of Engineering, King’s College London(工程学院,伦敦国王学院)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.AI

Comments Accepted to IEEE Robotics and Automation Letters (RAL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05925 2025-09-09 cs.CV cs.IT math.IT 88%

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models

Ruiqi Shen, Haotian Wu, Wenjing Zhang, Jiangjing Hu, Deniz Gunduz

机构 * Department of Electrical and Electronic Engineering, Imperial College London(帝国理工学院电子与电气工程系)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

Comments Published as a conference paper at IEEE 35th Workshop on Machine Learning for Signal Processing (MLSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23502 2025-07-15 cs.CV 88%

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching

Mengxiao Tian, Xinxiao Wu, Shuo Yang

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China(北京智能信息技术重点实验室,计算机科学与技术学院,北京理工大学,中国) Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, China(广东机器感知与智能计算实验室,深圳MSU-BIT大学,中国) Beijing Research Center of Intelligent Equipment for Agriculture, China(北京智能农业设备研究中心,中国)

专题命中 图文多模态 :multi-modal(title,abstract);image-text(title,abstract);分类 cs.CV

Comments accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17810 2025-04-11 cs.CV 88%

EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning

Yaxiong Wang, Yujiao Wu, Lianwei Wu, Lechao Cheng, Zhun Zhong, Meng Wang

专题命中 图文多模态 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11795 2025-04-08 cs.CV 88%

EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model

Feipeng Ma, Yizhou Zhou, Zheyu Zhang, Shilin Yan, Hebei Li, Zilong He, Siying Wu, Fengyun Rao, Yueyi Zhang, Xiaoyan Sun

专题命中 图文多模态 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12346 2024-09-26 cs.CV cs.IR cs.LG 88%

Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval

Naoya Sogi, Takashi Shibata, Makoto Terao

专题命中 图文多模态 :cross-modal(title,abstract);image-text(title,abstract);分类 cs.CV

Comments ECCV 2024. Code: https://github.com/NEC-N-SOGI/query-perturbation

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14842 2024-08-28 cs.CV cs.LG 88%

From Bias to Balance: Detecting Facial Expression Recognition Biases in Large Multimodal Foundation Models

Kaylee Chhua, Zhoujinyi Wen, Vedant Hathalia, Kevin Zhu, Sean O'Brien

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16510 2024-03-14 cs.CV 88%

Source-Free Domain Adaptation with Frozen Multimodal Foundation Model

Song Tang, Wenxin Su, Mao Ye, Xiatian Zhu

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08044 2024-01-22 cs.CV 88%

Benchmarking Robustness of Multimodal Image-Text Models under Distribution Shift

Jielin Qiu, Yi Zhu, Xingjian Shi, Florian Wenzel, Zhiqiang Tang, Ding Zhao, Bo Li, Mu Li

专题命中 图文多模态 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CV

Comments Accepted by Journal of Data-centric Machine Learning Research (DMLR) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01064 2023-11-03 cs.CV cs.LG 88%

Multimodal Foundation Models for Zero-shot Animal Species Recognition in Camera Trap Images

Zalan Fabian, Zhongqi Miao, Chunyuan Li, Yuanhan Zhang, Ziwei Liu, Andrés Hernández, Andrés Montes-Rojas, Rafael Escucha, Laura Siabatto, Andrés Link, Pablo Arbeláez, Rahul Dodhia, Juan Lavista Ferres

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

Comments 18 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏