arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2406.04309 2024-06-07 cs.CV cs.GR cs.LG cs.MM 76%

ReFiNe: Recursive Field Networks for Cross-modal Multi-scene Representation

Sergey Zakharov, Katherine Liu, Adrien Gaidon, Rares Ambrus

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV、cs.MM

Comments SIGGRAPH 2024. Project Page: https://zakharos.github.io/projects/refine/

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00281 2024-01-31 cs.CV cs.CL 76%

Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models

Zhuowan Li, Cihang Xie, Benjamin Van Durme, Alan Yuille

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.CL

Comments Accepted to EACL 2024. Code is released at https://github.com/Lizw14/visual_probing

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.05833 2023-10-18 cs.CV cs.HC cs.MM 76%

COLD Fusion: Calibrated and Ordinal Latent Distribution Fusion for Uncertainty-Aware Multimodal Emotion Recognition

Mani Kumar Tellamekala, Shahin Amiriparian, Björn W. Schuller, Elisabeth André, Timo Giesbrecht, Michel Valstar

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.MM

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11993 2023-04-26 cs.CV cs.MM 76%

MMC: Multi-Modal Colorization of Images using Textual Descriptions

Subhankar Ghosh, Saumik Bhattacharya, Prasun Roy, Umapada Pal, Michael Blumenstein

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.MM

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00495 2023-04-04 cs.CV cs.AI 76%

Multimodal Hyperspectral Image Classification via Interconnected Fusion

Lu Huo, Jiahao Xia, Leijie Zhang, Haimin Zhang, Min Xu

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

Comments 11 pages, five figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05163 2023-02-17 cs.MM cs.SD eess.AS 76%

Multimodal Dyadic Impression Recognition via Listener Adaptive Cross-Domain Fusion

Yuanchao Li, Peter Bell, Catherine Lai

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.MM、eess.AS

Comments Accepted to ICASSP2023. arXiv admin note: substantial text overlap with arXiv:2203.13932

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07289 2022-11-15 cs.CV cs.CL cs.LG 76%

Learning to Model Multimodal Semantic Alignment for Story Visualization

Bowen Li, Thomas Lukasiewicz

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.CL

Comments EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.05253 2022-10-26 cs.CV cs.CL 76%

MAGMA -- Multimodal Augmentation of Generative Models through Adapter-based Finetuning

Constantin Eichenberg, Sidney Black, Samuel Weinbach, Letitia Parcalabescu, Anette Frank

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.CL

Comments 13 pages, 6 figures, 2 tables. Minor improvements. Accepted at EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.00048 2022-03-29 cs.CV cs.AI 76%

Multi-modal Alignment using Representation Codebook

Jiali Duan, Liqun Chen, Son Tran, Jinyu Yang, Yi Xu, Belinda Zeng, Trishul Chilimbi

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments Accepted by CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.10245 2021-11-22 cs.LG cs.AI cs.CV 76%

Ubi-SleepNet: Advanced Multimodal Fusion Techniques for Three-stage Sleep Classification Using Ubiquitous Sensing

Bing Zhai, Yu Guan, Michael Catt, Thomas Ploetz

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

Comments Accepted in IMWUT for 2021 Dec issue

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.15124 2021-06-01 cs.CL cs.CV 76%

Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTs

Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, Desmond Elliott

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.CL

Comments To appear in TACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01394 2021-04-06 cs.CV cs.CL cs.LG 76%

MMBERT: Multimodal BERT Pretraining for Improved Medical VQA

Yash Khare, Viraj Bagal, Minesh Mathew, Adithi Devi, U Deva Priyakumar, CV Jawahar

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.11691 2019-12-30 cs.CV cs.MM 76%

Multi-Modal Attention-based Fusion Model for Semantic Segmentation of RGB-Depth Images

Fahimeh Fooladgar, Shohreh Kasaei

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.05733 2019-11-20 cs.RO cs.AI cs.CV cs.LG 76%

Choosing Smartly: Adaptive Multimodal Fusion for Object Detection in Changing Environments

Oier Mees, Andreas Eitel, Wolfram Burgard

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

Comments Published at the 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems. Added a new baseline with respect to the IROS version. Project page with code, pretrained models and our InOutDoorPeople RGB-D dataset at http://adaptivefusion.cs.uni-freiburg.de/

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.03678 2019-11-12 cs.CL cs.CV 76%

Bootstrapping Disjoint Datasets for Multilingual Multimodal Representation Learning

Ákos Kádár, Grzegorz Chrupała, Afra Alishahi, Desmond Elliott

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05683 2026-08-07 cs.CV cs.AI cs.MM 新提交 75%

DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation

DistMedVL:面向不确定性感知医学图像分割的分布视觉-语言对齐方法

Jiaxuan Li, Qing Xu, Xiangjian He, Yue Li, Daokun Zhang, Fiseha B. Tesema, Rong Qu

机构 * University of Nottingham Ningbo China(宁波诺丁汉大学) University of Nottingham(诺丁汉大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 针对医学图像分割中视觉-语言跨模态对齐的不确定性问题,提出DistMedVL概率框架,含PCM-Adapter等模块,在8个基准上以6.3M参数实现优于SOTA的性能,数据效率等更优。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09522 2026-05-14 cs.CV cs.AI cs.CL 75%

Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding

重新审视你所见:揭示视觉语义以引导LVLM解码

Beomsik Cho, Jaehyung Kim

机构 * Yonsei University(延世大学)

专题命中 多模态训练与对齐 :multimodal(abstract,abstract_cn);分类 cs.CV、cs.CL、cs.AI

AI总结 本文通过分析发现视觉token在幻觉中仍提供有效信息,提出ReVisiT方法通过参考视觉token引导LVLM解码,减少计算成本并提升性能。

Comments ACL 2026 Main Conference (Oral). 30 pages, 10 figures. Code: https://github.com/bscho333/ReVisiT

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22856 2026-02-02 cs.LG 75%

OptiMAG: Structure-Semantic Alignment via Unbalanced Optimal Transport

OptiMAG: 通过不平衡最优传输实现结构-语义对齐

Yilong Zuo, Xunkai Li, Zhihan Zhang, Qiangqiang Dai, Ronghua Li, Guoren Wang

机构 * Beijing Institute of Technology, Beijing, China(北京理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract)

AI总结 OptiMAG通过不平衡最优传输解决多模态图中结构与语义不一致问题,提升节点表示学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15272 2025-12-03 cs.CV cs.AI cs.CL 75%

CT-GLIP: 3D Grounded Language-Image Pretraining with CT Scans and Radiology Reports for Full-Body Scenarios

CT-GLIP:基于CT扫描和放射报告的3D grounded语言-图像预训练,用于全身场景

Jingyang Lin, Yingda Xia, Jianpeng Zhang, Ke Yan, Kai Cao, Le Lu, Jiebo Luo, Ling Zhang

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) University of Rochester(罗切斯特大学) Hupan Lab, \postcode 10587, \state Hangzhou, \country China(华普实验室,浙江省杭州市,中国) Department of Radiology, Shanghai Institution of Pancreatic Disease, \state Shanghai, \country China(胰腺疾病研究所放射科,上海市,中国)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 CT-GLIP通过构建细粒度CT报告对,提升3D grounded语言-图像预训练,实现更精确的跨模态对齐,从而在零样本任务中提升器官识别和肿瘤检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12609 2025-11-25 cs.CL cs.AI cs.CV 75%

Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data

Uni-MoE-2.0-Omni: 通过先进MoE、训练和数据扩展语言导向的多模态大模型

Yunxin Li, Xinyu Chen, Shenyuan Jiang, Haoyuan Shi, Zhenyu Liu, Xuanyu Zhang, Nanhao Deng, Zhenran Xu, Yicheng Ma, Meishan Zhang, Baotian Hu, Min Zhang

机构 * Research Institute of Computing and Intelligence(计算与智能研究 institute) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 Uni-MoE-2.0-Omni通过先进MoE、训练和数据技术,实现了语言导向的多模态大模型,展现出在多模态理解、推理和生成任务中的卓越性能。

Comments 47 pages,10 Figures, Project Website: https://idealistxy.github.io/Uni-MoE-v2.github.io/ Codes: https://github.com/HITsz-TMG/Uni-MoE

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08645 2025-10-31 cs.LG 75%

When Kernels Multiply, Clusters Unify: Fusing Embeddings with the Kronecker Product

Youqi Wu, Jingwei Zhang, Farzan Farnia

机构 * Department of Computer Science & Engineering, The Chinese University of Hong Kong(计算机科学与工程系,香港中文大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10666 2025-10-07 astro-ph.SR astro-ph.GA astro-ph.IM 75%

Machine-learning inference of stellar properties using integrated photometric and spectroscopic data

Ilay Kamai, Alex M. Bronstein, Hagai B. Perets

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract)

Comments Accepted to ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26625 2025-10-01 cs.LG cs.AI cs.CV cs.MM 75%

Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

Junlin Han, Shengbang Tong, David Fan, Yufan Ren, Koustuv Sinha, Philip Torr, Filippos Kokkinos

机构 * Meta Superintelligence Labs(Meta 超智能实验室) University of Oxford(牛津大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Project page: https://junlinhan.github.io/projects/lsbs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12363 2025-09-10 cs.CV cs.AI cs.CL cs.LG cs.RO 75%

Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts

Qi Feng

机构 * Kyoto University(京都大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 26 pages, 19 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20109 2025-03-24 cs.CV cs.AI cs.MM 75%

GiVE: Guiding Visual Encoder to Perceive Overlooked Information

Junjie Li, Jianghong Ma, Xiaofeng Zhang, Yuhang Li, Jianyang Shi

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI、cs.MM

Comments This paper was accepted by ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01483 2024-11-18 cs.CV cs.AI cs.CL 75%

MANTIS: Interleaved Multi-Image Instruction Tuning

Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, Wenhu Chen

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 13 pages, 3 figures, 13 tables

Journal ref Transactions on Machine Learning Research 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20178 2024-11-13 cs.AI cs.CL cs.CV cs.LG 75%

LLMs Can Evolve Continually on Modality for X-Modal Reasoning

Jiazuo Yu, Haomiao Xiong, Lu Zhang, Haiwen Diao, Yunzhi Zhuge, Lanqing Hong, Dong Wang, Huchuan Lu, You He, Long Chen

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17827 2024-11-12 cs.CV cs.AI cs.CL cs.LG 75%

Unified Lexical Representation for Interpretable Visual-Language Alignment

Yifan Li, Yikai Wang, Yanwei Fu, Dongyu Ru, Zheng Zhang, Tong He

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13173 2024-07-17 cs.CV cs.AI cs.CL cs.LG 75%

Biomedical Visual Instruction Tuning with Clinician Preference Alignment

Hejie Cui, Lingjun Mao, Xin Liang, Jieyu Zhang, Hui Ren, Quanzheng Li, Xiang Li, Carl Yang

专题命中 多模态训练与对齐 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18765 2024-03-14 cs.CV cs.AI cs.CL cs.LG 75%

MLLMs-Augmented Visual-Language Representation Learning

Yanqing Liu, Kai Wang, Wenqi Shao, Ping Luo, Yu Qiao, Mike Zheng Shou, Kaipeng Zhang, Yang You

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏