arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2301.11362 2023-01-30 cs.CV cs.CL 84%

Improving Cross-modal Alignment for Text-Guided Image Inpainting

Yucheng Zhou, Guodong Long

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments EACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00678 2022-12-02 cs.CL cs.CV cs.LG 84%

Adapted Multimodal BERT with Layer-wise Fusion for Sentiment Analysis

Odysseas S. Chlapanis, Georgios Paraskevopoulos, Alexandros Potamianos

专题命中 多模态训练与对齐 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.10916 2022-11-29 cs.CV cs.CL cs.LG 84%

Hierachical Delta-Attention Method for Multimodal Fusion

Kunjal Panchal

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Need to update the results

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00526 2022-11-02 cs.CL cs.AI 84%

Leveraging Graph-based Cross-modal Information Fusion for Neural Sign Language Translation

Jiangbin Zheng, Siyuan Li, Cheng Tan, Chong Wu, Yidong Chen, Stan Z. Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09227 2022-04-21 cs.CL cs.SD eess.AS 84%

Cross-stitched Multi-modal Encoders

Karan Singla, Daniel Pressel, Ryan Price, Bhargav Srinivas Chinnari, Yeon-Jun Kim, Srinivas Bangalore

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05784 2022-03-14 eess.IV cs.AI cs.CV 84%

AI-enabled Automatic Multimodal Fusion of Cone-Beam CT and Intraoral Scans for Intelligent 3D Tooth-Bone Reconstruction and Clinical Applications

Jin Hao, Jiaxiang Liu, Jin Li, Wei Pan, Ruizhe Chen, Huimin Xiong, Kaiwei Sun, Hangzheng Lin, Wanlu Liu, Wanghui Ding, Jianfei Yang, Haoji Hu, Yueling Zhang, Yang Feng, Zeyu Zhao, Huikai Wu, Youyi Zheng, Bing Fang, Zuozhu Liu, Zhihe Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 30 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.14457 2022-01-06 cs.CL cs.AI cs.LG 84%

Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning

Subhojeet Pramanik, Shashank Mujumdar, Hima Patel

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.13731 2021-08-11 cs.CV cs.AI 84%

UIBert: Learning Generic Multimodal Representations for UI Understanding

Chongyang Bai, Xiaoxue Zang, Ying Xu, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, Blaise Aguera y Arcas

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 8 pages, IJCAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.11672 2021-05-26 cs.CL cs.CV cs.LG 84%

ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents

Weihong Lin, Qifang Gao, Lei Sun, Zhuoyao Zhong, Kai Hu, Qin Ren, Qiang Huo

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments To be published at ICDAR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03435 2021-04-09 cs.CV cs.AI 84%

Multimodal Fusion Refiner Networks

Sethuraman Sankaran, David Yang, Ser-Nam Lim

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.04727 2021-01-14 cs.CL cs.AI 84%

Latent Alignment of Procedural Concepts in Multimodal Recipes

Hossein Rajaby Faghihi, Roshanak Mirzaee, Sudarshan Paliwal, Parisa Kordjamshidi

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments Published in ALVR 2020, a workshop in ACL 2020

Journal ref Proceedings of the First Workshop on Advances in Language and Vision Research 2020 (26-31)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.00514 2020-10-02 cs.CV cs.CL 84%

Referring Image Segmentation via Cross-Modal Progressive Comprehension

Shaofei Huang, Tianrui Hui, Si Liu, Guanbin Li, Yunchao Wei, Jizhong Han, Luoqi Liu, Bo Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by CVPR 2020. Code is available at https://github.com/spyflying/CMPC-Refseg

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.11740 2020-07-21 cs.CV cs.CL cs.LG 84%

UNITER: UNiversal Image-TExt Representation Learning

Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, Jingjing Liu

专题命中 多模态训练与对齐 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.07407 2018-11-20 cs.CV cs.AI cs.LG 84%

Multimodal Densenet

Faisal Mahmood, Ziyun Yang, Thomas Ashley, Nicholas J. Durr

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.03414 2018-10-09 cs.CV cs.MM 84%

Dense Multimodal Fusion for Hierarchically Joint Representation

Di Hu, Feiping Nie, Xuelong Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.03920 2018-08-14 cs.LG cs.AI cs.CL cs.NE stat.ML 84%

Multimodal Language Analysis with Recurrent Multistage Fusion

Paul Pu Liang, Ziyin Liu, Amir Zadeh, Louis-Philippe Morency

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.09534 2016-11-30 cs.CV cs.CL 84%

Is a picture worth a thousand words? A Deep Multi-Modal Fusion Architecture for Product Classification in e-commerce

Tom Zahavy, Alessandro Magnani, Abhinandan Krishnan, Shie Mannor

专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20198 2026-02-03 cs.CV 84%

A Survey of Token Compression for Efficient Multimodal Large Language Models

多模态大语言模型高效性中的标记压缩综述

Kele Shao, Keda Tao, Kejia Zhang, Sicheng Feng, Mu Cai, Yuzhang Shang, Haoxuan You, Can Qin, Yang Sui, Huan Wang

机构 * Zhejiang University(浙江大学) Westlake University(西湖大学) Xiamen University(厦门大学) National University of Singapore(新加坡国立大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Central Florida(佛罗里达大学) Salesforce AI Research(Salesforce AI研究) Rice University(德克萨斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文综述了多模态大语言模型中标记压缩技术,分类讨论了图像、视频和音频三种模态的压缩方法及其机制,旨在推动该领域的发展。

Comments For ongoing updates and to track the latest advances in this promising area, we maintain a public repository: https://github.com/cokeshao/Awesome-Multimodal-Token-Compression

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19491 2024-07-30 cs.CV 84%

Multi-modal Crowd Counting via Modal Emulation

Chenhao Wang, Xiaopeng Hong, Zhiheng Ma, Yupeng Wei, Yabin Wang, Xiaopeng Fan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments This is the preprint version of the paper to appear in BMVC 2024. Please cite the final published version. Code is available at https://github.com/Mr-Monday/Multi-modal-Crowd-Counting-via-Modal-Emulation

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09513 2024-03-15 cs.CR cs.AI 84%

AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting

Yu Wang, Xiaogeng Liu, Yu Li, Muhao Chen, Chaowei Xiao

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Multimodal Large Language Models Defense, 25 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12898 2023-08-28 cs.MM cs.AI cs.CL cs.CV 84%

Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?

Fei Wang, Liang Ding, Jun Rao, Ye Liu, Li Shen, Changxing Ding

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments [TL;DR] we design and release the SNARE, the first large-scale multimodal alignment probing benchmark for current vision-language pretrained models

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11991 2026-05-04 cs.CV cs.AI cs.CL 83%

VGR: Visual Grounded Reasoning

VGR:视觉基础推理

Jiacong Wang, Zijian Kang, Haochen Wang, Haiyong Jiang, Jiawen Li, Bohong Wu, Ya Wang, Jiao Ran, Xiao Liang, Chao Feng, Jun Xiao

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ByteDance Inc.(字节跳动公司)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出VGR,一种增强视觉感知的多模态大语言模型,通过图像区域检测与回放提升多模态推理能力,在多个基准测试中表现优异。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14152 2026-08-17 cs.AI 新提交 83%

Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach

面向科技情报(STI)的高效多模态多语言观点抽取:一种基于QLoRA的微调方法

Sheng Hong, Xuanqi Wang, Jiacheng Wang, Yuwei Wang

机构 * Beihang University(北京航空航天大学) Nanchang University(南昌大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

AI总结 本研究针对科技情报观点抽取的噪声过滤与结构化输出问题,提出基于QLoRA微调的多模态框架,在2194样本数据集上实现多语言观点抽取性能显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13973 2026-08-17 cs.CV 新提交 83%

Rethinking Auxiliary Modalities in Multimodal Zero-shot Anomaly Detection: From Semantic Fusion to Conditional Modulation

重新思考多模态零样本异常检测中的辅助模态:从语义融合到条件调制

Peng Wu, Xin Ge, Yujia Sun, Guansong Pang

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 本研究针对现有多模态零样本异常检测方法的缺陷,提出即插即用的辅助条件增强框架,通过全局到局部的条件调制实现选择性多模态增强,在MVTec 3D-AD等数据集上提升了现有RGB零样本异常检测器的性能并达到最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13741 2026-08-17 cs.CL cs.LG 新提交 83%

GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis

GALA:面向文本到时间序列合成的生成感知跨模态对齐

Haochen Zhang, Gengwei Zhang, Laura Yao, Nicholas Knoz, Tianlong Chen

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CL

AI总结 本研究针对文本到时间序列合成中条件表示与信号模态不匹配的问题,提出GALA两阶段跨模态对齐方法,在TSFragment-600K数据集上实现SOTA,打破了生成器内部文本编码器的保真度与贴合度权衡。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13385 2026-08-14 cs.CV 新提交 83%

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL

何时仅需任务向量?隐式多模态上下文学习的经验理论

Jiaqian Li

机构 * Brown University(布朗大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出选择-实现假说,通过受控多模态任务和VQA基准研究发现,静态任务向量的成功取决于演示诱导变化的跨查询共享程度,为隐式多模态上下文学习的方法选择提供了统一经验理论。

Comments Accepted by Empirical Theory in Representation Learning @ ECCV 2026, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11685 2026-08-13 cs.CV 新提交 83%

EGM-Det: Entropy-Guided Multimodal Adaptive Fusion for UAV RGB-IR Object Detection

EGM-Det:面向无人机RGB-IR目标检测的熵引导多模态自适应融合

Cunzheng Fan, Dawei Yan, Guanlin Wang, Xingshuo Yang, Yupeng Jia, Jing Yang, Haokui Zhang

机构 * School of Cybersecurity, Northwestern Polytechnical University(西北工业大学网络空间安全学院) School of Automation and Software Engineering, Shanxi University(山西大学自动化与软件学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出EGM-Det框架,通过熵引导多模态自适应融合解决无人机RGB-IR目标检测中模态可靠性被忽略的问题,在三个数据集上实现最优性能,VEDAI上较现有方法提升超10个百分点。

Comments 14 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07917 2026-08-12 cs.AI 版本更新 83%

TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

TongGuOCR:面向中文历史文献的布局感知与 token 增强 OCR 框架

Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin

机构 * School of Electronic and Information Engineering, South China University of Technology(华南理工大学电子与信息工程学院) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 针对中文历史文献OCR的复杂布局、生僻字等挑战,提出TongGuOCR框架,通过布局感知预处理与token增强识别模块,在M5HisDoc等基准上取得优于同类模型的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15779 2026-08-12 cs.CV 版本更新 83%

Multimodal Ambivalence and Hesitancy Recognition via Cross-Attention and Gated Fusion

通过交叉注意力和门控融合进行多模态矛盾与犹豫识别

Oussama Berhili, Yassine Ouzar, Larbi Boubchir

机构 * University of Paris 8(巴黎第八大学) LIASD Laboratory(LIASD实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 为ECCV 2026的ABAW11挑战赛开发多模态框架,用预训练编码器提取多模态特征,建立单模态基线,在此基础上提出含交叉注意力和门控融合的多模态架构,验证集宏F1达0.7394,提升显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23755 2026-08-11 cs.CV cs.RO 交叉投稿 83%

DAP-Pose: Deep Temporal Alignment and Physics-aware Cross-modal Sensor Fusion for Robust Pose Estimation

DAP-Pose:用于鲁棒姿态估计的深度时间对齐和物理感知跨模态传感器融合

Jianhan Lin, Yuchu Qin, Jiateng Yuan, Wenbo Zhang, Shuai Gao

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院) International Research Center of Big Data for Sustainable Development Goals(可持续发展大数据国际研究中心) University of Chinese Academy of Sciences(中国科学院大学) The University of Adelaide(阿德莱德大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 针对复杂环境下多模态传感器的姿态估计问题,提出DAP-Pose模型,通过双级跨模态融合模块捕捉线索,深度时间对齐模块处理异步流,结合物理感知约束,在KITTI数据集上达最优性能,平均平移误差1.31%,旋转误差0.46°,严重错位下也能准确估计。

详情

展开后加载摘要…

URL PDF HTML 收藏