arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45787 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4633 篇

2311.07766 2023-11-15 cs.CV cs.AI cs.CL cs.LG 87%

Vision-Language Integration in Multimodal Video Transformers (Partially) Aligns with the Brain

Dota Tianai Dong, Mariya Toneva

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.00621 2023-06-13 cs.CL cs.AI cs.CV cs.LG 87%

Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training

Yan Zeng, Wangchunshu Zhou, Ao Luo, Ziming Cheng, Xinsong Zhang

专题命中 图文多模态 :cross-modal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15016 2023-03-28 cs.CL cs.AI cs.IR cs.MM 87%

Borrowing Human Senses: Comment-Aware Self-Training for Social Media Multimodal Classification

Chunpu Xu, Jing Li

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CL、cs.AI、cs.MM

Comments accepted to EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17806 2026-07-21 cs.AI 新提交 86%

PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model

PGN:基于盘古多模态基础模型的视觉语言导航系统的设计与实现

Li Xian, Mingxi Li, Yizheng Wang, Yiming Shen, Qi Chen, Zhuoling Xiao

机构 * University of Electronic Science and Technology of China(电子科技大学) School of Information and Communication Engineering(信息与通信工程学院)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.AI

AI总结 研究基于盘古多模态基础模型设计实现视觉语言导航系统PGN,训练分两阶段,先对齐视觉与语言编码器,再使模型适应专家轨迹,结合多种计算方式,在特定评估下取得一定指标,量化了离线专家动作对齐情况。

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06646 2026-02-02 cs.CV cs.LG 86%

The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models

窄门:原生多模态模型中的局部图像-文本通信

Alessandro Pietro Serra, Francesco Ortu, Emanuele Panizon, Lucrezia Valeriani, Lorenzo Basile, Alessio Ansuini, Diego Doimo, Alberto Cazzaniga

机构 * Area Science Park, Trieste, Italy(特里斯特科学公园) SISSA, Trieste, Italy(SISSA) University of Trieste, Trieste, Italy(特里斯特大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(title);分类 cs.CV

AI总结 研究揭示了原生多模态模型中图像与文本信息交互的机制,发现通过单个标记作为窄门影响图像理解性能,提出基于标记级干预的精细控制方法。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11177 2025-11-26 cs.CV 86%

Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation

Viper-F1:基于交叉模态状态空间调制的高效细粒度多模态理解

Quoc-Huy Trinh

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(title);分类 cs.CV

AI总结 Viper-F1通过引入高效液态状态空间动力学和令牌-网格相关模块,实现了高效细粒度多模态理解。

Comments arXiv admin comment: This version has been removed by arXiv administrators as the submitter did not have the rights to agree to the license at the time of submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11705 2025-11-18 cs.LG cs.CV 86%

Multimodal ML: Quantifying the Improvement of Calorie Estimation Through Image-Text Pairs

Arya Narang

专题命中 图文多模态 :multimodal(title,abstract);image-text(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12077 2024-12-17 cs.CV 86%

CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology

Yuxuan Sun, Yixuan Si, Chenglu Zhu, Xuan Gong, Kai Zhang, Pingyi Chen, Ye Zhang, Zhongyi Shui, Tao Lin, Lin Yang

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.CV

Comments 22 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07127 2023-04-17 cs.CL 86%

OPI at SemEval 2023 Task 1: Image-Text Embeddings and Multimodal Information Retrieval for Visual Word Sense Disambiguation

Sławomir Dadas

专题命中 图文多模态 :multimodal(title,abstract);image-text(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09099 2022-04-21 cs.CL 86%

AMS_ADRN at SemEval-2022 Task 5: A Suitable Image-text Multimodal Joint Modeling Method for Multi-task Misogyny Identification

Da Li, Ming Yi, Yukai He

专题命中 图文多模态 :multimodal(title,abstract);image-text(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.07966 2020-01-24 cs.CV 86%

ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, Arun Sacheti

专题命中 图文多模态 :image-text(title,abstract);cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28640 2026-08-03 cs.CL cs.CV 新提交 86%

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

TokenSwap:多模态大语言模型中模态差距的基准测试与缩减

Andong Hua, Colton Bishop, Igor Mordatch, Arian Hosseini, Jindong Gu, Aleksandra Faust, Rebecca Roelofs, Yao Qin

机构 * University of California, Santa Barbara(加州大学圣巴巴拉分校) Google DeepMind(谷歌DeepMind)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出TokenSwap方法构建基准TokenSwap-Bench,发现42个多模态大语言模型普遍存在模态差距,推理模型差距更小,训练中引入TokenSwap可有效缓解该差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17588 2026-06-16 cs.CV cs.CL 版本更新 86%

Dual-branch Prompting for Multimodal Machine Translation

双分支提示用于多模态机器翻译

Jie Wang, Zhendong Yang, Liansong Zong, Xiaobo Zhang, Dexian Wang, Ji Zhang

机构 * School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院) School of Computer and Software Engineering, Xihua University(西华大学计算机与软件工程学院) School of Intelligent Medicine, Chengdu University of Traditional Chinese Medicine(成都中医药大学针灸推拿学院)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

AI总结 提出基于扩散模型的双分支提示框架D2P-MMT,利用重建图像过滤视觉噪声,通过分布对齐损失提升鲁棒翻译性能。

Comments This manuscript has been fully accepted and published by ACM Transactions on Multimedia Computing, Communications, and Applications (ACM TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00105 2026-06-02 cs.CV cs.AI 86%

Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning

视觉噪声引导的上下文蒸馏用于多模态大语言模型遗忘

Junkai Chen, Yuhao He, Junxiang You, Ruiqi Liu, Chenyu Wang, Shu Wu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Advanced Interdisciplinary Sciences, UCAS(北京大学交叉学科研究院)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 提出视觉噪声引导的上下文蒸馏(VGID)框架,通过双模态干预构建教师分布进行蒸馏,实现多模态大语言模型参数级遗忘,平衡遗忘效果与模型效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27916 2026-05-28 cs.CV cs.CL 86%

OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models

OphIn-500K:策划网络规模的视觉指令以扩展眼科多模态大语言模型

Xuanzhao Dong, Wenhui Zhu, Xiwen Chen, Hao Wang, Xin Li, Yujian Xiong, Jiajun Cheng, Jingjing Wang, Xiaobing Yu, Haiyu Wu, Shao Tang, Zhipeng Wang, Langechuan Liu, Shan Lin, Oana Dumitrascu, Yalin Wang

机构 * Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学) Washington University in St. Louis(圣路易斯华盛顿大学) University of Notre Dame(诺特丹大学) Florida State University(佛罗里达州立大学) Rice University(里德大学) NVIDIA(英伟达) Mayo Clinic(梅奥诊所)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.CL

AI总结 提出OphIn-Engine流水线从网络视频中构建高质量眼科指令数据,生成包含50万+指令实例的OphIn-500K数据集,并基于此开发眼科专用多模态大语言模型OphIn-VL,在多项任务上超越现有通用医学和专用模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07574 2026-05-28 cs.CV cs.CL 86%

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention

ViCA:仅视觉交叉注意力的高效多模态大语言模型

Wenjie Liu, Hao Wu, Xin Qiu, Xudong Wang, Yingqi Fan, Yihan Zhang, Anhao Zhao, Yunpu Ma, Xiaoyu Shen

机构 * Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究院、东部技术研究院) Munich Center for Machine Learning, LMU Munich(慕尼黑机器学习中心、慕尼黑大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.CL

AI总结 提出ViCA架构,通过仅视觉交叉注意力减少视觉令牌计算,在保持98%准确率的同时将视觉计算降至4%,实现显著加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07394 2026-05-11 cs.CV cs.AI 86%

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning

BalCapRL:一种基于RL的MLLM图像描述的平衡框架

Shaokai Ye, Vasileios Saveris, Yihao Qian, Jiaming Hu, Elmira Amirloo, Peter Grasch

机构 * Apple(苹果公司)

专题命中 图文多模态 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出BalCapRL框架,通过联合优化正确性、参考覆盖和语言质量,解决图像描述中因评价指标狭窄导致的权衡问题,并通过改进的奖励解耦归一化和长度条件奖励遮蔽提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00733 2026-05-04 cs.NI cs.AI cs.LG cs.MM 86%

EASE: Federated Multimodal Unlearning via Entanglement-Aware Anchor Closure

EASE: 通过纠缠感知锚闭合实现联邦多模态去学习

Zihao Ding, Beining Wu, Jun Huang

机构 * Department of Electrical Engineering and Computer Science(电气工程与计算机科学系)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.AI、cs.MM

AI总结 本文提出EASE框架,通过闭合三个锚通道解决联邦多模态学习中的遗忘知识纠缠问题,采用余弦-正弦分解和方向选择性遗忘锁等方法,在多个数据集上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06090 2026-03-09 cs.CV cs.CL 86%

DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model

DeepSight: 通过深度驱动的多模态模型连接深度图与语言

Hao Yang, Hongbo Zhang, Yanyan Zhao, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV、cs.CL

AI总结 DeepSight通过深度驱动的多模态模型提升三维场景理解,利用深度图像特性增强空间推理能力,并通过新数据集和模型改进实现了深度感知和任务性能的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09396 2026-03-05 cs.CL cs.AI 86%

Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque

多模态大语言模型用于低资源语言:巴斯克语的案例研究

Lukas Arana, Julen Etxaniz, Ander Salaberria, Gorka Azkune

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CL、cs.AI

AI总结 本文通过开发巴斯克语多模态大语言模型,证明低比例多模态数据即可获得良好性能,且无需巴斯克语指导的LLM即可实现强模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04587 2026-02-23 cs.CL cs.AI cs.CY 86%

VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration

VILLAIN 在 AVerImaTeC 上:通过多智能体协作验证图像-文本声明

Jaeyoon Jung, Yejun Yoon, Kunwoo Park

机构 * School of AI Convergence, Soongsil University(人工智能融合学院,顺世大学) MAUM AI Inc.(MAUM人工智能公司) Department of Intelligent Semiconductors, Soongsil University(智能半导体系,顺世大学)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

AI总结 VILLAIN 通过多智能体协作验证图像-文本声明,实现事实核查系统的高效准确验证。

Comments A system description paper for the AVerImaTeC shared task at the Ninth FEVER Workshop (co-located with EACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13758 2026-02-17 cs.CV cs.AI 86%

OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding

OmniScience: 一个大规模多模态数据集用于科学图像理解

Haoyi Tao, Chaozheng Huang, Nan Wang, Han Lyu, Linfeng Zhang, Guolin Ke, Xi Fang

机构 * DP Technology(DP技术)

专题命中 图文多模态 :multi-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 OmniScience是一个大规模多模态数据集,通过动态模型路由生成高信息密度的图像标题,提升多模态模型在科学图像理解上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09879 2026-01-16 cs.CV cs.AI 86%

MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation

MedVL-SAM2:一种统一的3D医学视觉-语言模型,用于多模态推理和基于提示的分割

Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong

机构 * Department of Biomedical Engineering, University of Florida(佛罗里达大学生物医学工程系) Department of Radiology, University of Florida(佛罗里达大学放射学系) Research Computing, University of Florida(佛罗里达大学研究计算中心) Department of Medicine, University of Florida(佛罗里达大学医学系) Department of Radiology, UC San Francisco(旧金山大学放射学系)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 MedVL-SAM2是一种统一的3D医学多模态模型,通过联合训练实现报告生成、VQA和多任务分割的高性能表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20015 2025-12-02 cs.AI cs.CL 86%

Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak

通过多模态大语言模型 jailbreak 实现高效的 LLM-jailbreaking

Haoxuan Ji, Zheng Lin, Zhenxing Niu, Xinbo Gao, Gang Hua

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种通过多模态大语言模型实现高效 LLM-jailbreaking 的方法,结合图像-文本语义匹配提升攻击成功率,并在效率和泛化能力上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21705 2025-12-01 cs.CL cs.CV 86%

Insight-A: Attribution-aware for Multimodal Misinformation Detection

Insight-A: 多模态虚假信息检测中的归因意识

Junjie Wu, Yumeng Fu, Chen Gong, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 Insight-A通过归因意识和分层推理提升多模态虚假信息检测效果,有效识别伪造来源并增强跨模态一致性检查。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15162 2025-10-20 cs.CV cs.CL 86%

Train a Unified Multimodal Data Quality Classifier with Synthetic Data

Weizhi Wang, Rongmei Lin, Shiyang Li, Colin Lockard, Ritesh Sarkhel, Sanket Lokegaonkar, Jingbo Shang, Xifeng Yan, Nasser Zalmout, Xian Li

机构 * UC Santa Barbara(加州大学圣芭芭拉分校) Amazon Stores Foundational AI(亚马逊商店基础人工智能) UC San Diego(加州大学圣地亚哥分校)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06434 2025-09-26 cs.CV cs.AI 86%

CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment

Shengzhu Yang, Jiawei Du, Shuai Lu, Weihang Zhang, Ningli Wang, Huiqi Li

机构 * Beijing Institute of Technology(北京理工大学) Beijing Tongren Hospital(北京同仁医院)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20994 2025-07-29 cs.CV cs.AI 86%

Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM

Shen Li, Liuyi Yao, Wujia Niu, Lan Zhang, Yaliang Li

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Codes and data are available at https://github.com/listen0425/Security-Tensors

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08590 2025-07-14 cs.MM cs.CV 86%

Visual Semantic Description Generation with MLLMs for Image-Text Matching

Junyu Chen, Yihua Gao, Mingyong Li

机构 * Chongqing Normal University(重庆师范大学)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by ICME2025 oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23639 2025-07-01 cs.CV cs.AI 86%

Unified Multimodal Understanding via Byte-Pair Visual Encoding

Wanpeng Zhang, Yicheng Feng, Hao Luo, Yijiang Li, Zihao Yue, Sipeng Zheng, Zongqing Lu

机构 * Peking University(北京大学) UC San Diego(加州大学圣地亚哥分校) Renmin University of China(中国人民大学) BeingBeyond

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏