arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2310.00672 2023-10-10 cs.LG cs.CL cs.CV 73%

GeRA: Label-Efficient Geometrically Regularized Alignment

Dustin Klebe, Tal Shnitzer, Mikhail Yurochkin, Leonid Karlinsky, Justin Solomon

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10729 2023-09-28 cs.CV cs.AI cs.LG 73%

UnICLAM:Contrastive Representation Learning with Adversarial Masking for Unified and Interpretable Medical Vision Question Answering

Chenlu Zhan, Peng Peng, Hongsen Wang, Tao Chen, Hongwei Wang

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10354 2023-08-22 cs.AI cs.CL 73%

Imaginations of WALL-E : Reconstructing Experiences with an Imagination-Inspired Module for Advanced AI Systems

Zeinab Sadat Taghavi, Soroush Gooran, Seyed Arshan Dalili, Hamidreza Amirzadeh, Mohammad Jalal Nematbakhsh, Hossein Sameti

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments 18 pages,

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09695 2023-08-15 cs.CV cs.GR cs.MM 73%

PersonalTailor: Personalizing 2D Pattern Design from 3D Garment Point Clouds

Sauradip Nag, Anran Qi, Xiatian Zhu, Ariel Shamir

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15765 2023-05-26 cs.CV cs.AI 73%

Language-Guided 3D Object Detection in Point Cloud for Autonomous Driving

Wenhao Cheng, Junbo Yin, Wei Li, Ruigang Yang, Jianbing Shen

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09609 2023-04-20 cs.CV cs.AI 73%

MMDR: A Result Feature Fusion Object Detection Approach for Autonomous System

Wendong Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.02995 2023-03-07 cs.CV cs.CL cs.LG 73%

HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention

Shijie Geng, Jianbo Yuan, Yu Tian, Yuxiao Chen, Yongfeng Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00902 2023-02-06 cs.LG cs.CL cs.CV 73%

Language Quantized AutoEncoders: Towards Unsupervised Text-Image Alignment

Hao Liu, Wilson Yan, Pieter Abbeel

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Fixed typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07636 2022-12-06 cs.CV cs.CL cs.LG 73%

EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, Yue Cao

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments v2: (i) fix / update EVA IN-1K variants results. (ii) add / update EVA-CLIP results. (iii) add Appendix. (iv) release all the code and models at https://github.com/baaivision/EVA

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14777 2022-12-02 cs.CV cs.CL 73%

Alignment-Enriched Tuning for Patch-Level Pre-trained Document Image Models

Lei Wang, Jiabang He, Xing Xu, Ning Liu, Hui Liu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by AAAI 2023. Code is available at https://github.com/MAEHCM/AET

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06097 2022-11-14 cs.CV cs.AI 73%

Interactive Context-Aware Network for RGB-T Salient Object Detection

Yuxuan Wang, Feng Dong, Jinchao Zhu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.13979 2022-08-01 cs.CL cs.AI 73%

Knowing Where and What: Unified Word Block Pretraining for Document Understanding

Song Tao, Zijian Wang, Tiantian Fan, Canjie Luo, Can Huang

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments incomplete experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08387 2022-07-20 cs.CL cs.CV 73%

LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, Furu Wei

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments ACM Multimedia 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.15119 2022-05-26 cs.CV cs.AI 73%

Aerial Images Meet Crowdsourced Trajectories: A New Approach to Robust Road Extraction

Lingbo Liu, Zewei Yang, Guanbin Li, Kuo Wang, Tianshui Chen, Liang Lin

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments This work has been accepted by IEEE Transactions on Neural Networks and Learning Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.04482 2022-03-31 cs.CV cs.CL 73%

FLAVA: A Foundational Language And Vision Alignment Model

Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, Douwe Kiela

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13285 2022-03-30 cs.SD cs.CV cs.LG eess.AS 73%

Continuous-Time Audiovisual Fusion with Recurrence vs. Attention for In-The-Wild Affect Recognition

Vincent Karas, Mani Kumar Tellamekala, Adria Mallol-Ragolta, Michel Valstar, Björn W. Schuller

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、eess.AS

Comments 10 pages, 1 figures, added references and an overview figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14973 2021-12-23 cs.CV cs.AI cs.LG cs.RO 73%

MultiPath++: Efficient Information Fusion and Trajectory Aggregation for Behavior Prediction

Balakrishnan Varadarajan, Ahmed Hefny, Avikalp Srivastava, Khaled S. Refaat, Nigamaa Nayakanti, Andre Cornman, Kan Chen, Bertrand Douillard, Chi Pang Lam, Dragomir Anguelov, Benjamin Sapp

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.03331 2021-06-08 cs.CV cs.CL 73%

SelfDoc: Self-Supervised Document Representation Learning

Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, Hongfu Liu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments To appear in CVPR'2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.11562 2021-01-28 cs.CV cs.CL 73%

Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder Network

Yehao Li, Yingwei Pan, Ting Yao, Jingwen Chen, Tao Mei

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments AAAI 2021; Code is publicly available at: https://github.com/YehLi/TDEN

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.00200 2020-10-01 cs.CV cs.CL cs.LG 73%

HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, Jingjing Liu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.04464 2020-04-21 cs.CV cs.CL 73%

Relationship-Embedded Representation Learning for Grounding Referring Expressions

Sibei Yang, Guanbin Li, Yizhou Yu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments This paper is going to appear in TPAMI. Code is available at https://github.com/sibeiyang/sgmn/tree/master/lib/cmrin_models

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05491 2026-06-05 cs.CV cs.RO 72%

Unpaired RGB-Thermal Gaussian-Splatting Using Visual Geometric Transformers

无配对RGB-热成像高斯泼溅使用视觉几何变换器

Jean Cordonnier, Chenghao Xu, Olga Fink, Malcolm Mielle

机构 * Ecole Polytechnique Federale de Lausanne(瑞士联邦理工学院洛桑分校) Schindler EPFL Lab(施耐德EPFL实验室)

专题命中 多模态训练与对齐 :multi-modal(abstract,comments);cross-modal(abstract);分类 cs.CV

AI总结 提出一种无配对RGB-热成像新视角合成框架,利用VGGT估计各模态相机位姿并通过Procrustes对齐,结合多模态3D高斯泼溅实现联合重建,在保持RGB保真度的同时实现热成像视图合成。

Comments Accepted at ICRA 2026's Workshop MM-SpatialAI: Multi-Modal Spatial AI for Robust Navigation and Open-World Understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00086 2026-08-04 cs.CV cs.AI cs.CL cs.LG 版本更新 71%

Hierarchical Pre-Training of Vision Encoders with Large Language Model

基于大语言模型的视觉编码器分层预训练

Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai, Jiajie Diao, Chen-Yi Lee

机构 * University of Cincinnati(辛辛那提大学) National Yang Ming Chiao Tung University(国立阳明交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)

AI总结 本文提出HIVE框架,通过引入视觉编码器与大语言模型间的分层交叉注意力机制,提升视觉语言对齐,改进特征融合与表征学习,实验表明其在图像分类和多模态任务中表现优异。

Comments 17 pages, 14 figures, accepted to Computer Vision and Pattern Recognition Conference (CVPR) Workshops 2026. 5th MMFM Workshop: What is Next in Multimodal Foundation Models?

Journal ref In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7415-7424) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07407 2026-07-28 cs.LG 版本更新 71%

Emergent Symbolic Structure in Health Foundation Models: Extraction, Alignment, and Cross-Modal Transfer

健康基础模型中的涌现符号结构:提取、对齐与跨模态迁移

Gajendra Katuwal, Advait Koparkar, Salar Abbaspourazad, Anshuman Mishra, Sarvesh Kirthivasan

机构 * Apple(苹果公司)

专题命中 多模态训练与对齐 :cross-modal(title)

AI总结 本文提出一种训练后框架,通过分解冻结嵌入以提取可解释的符号,用于对齐嵌入空间。在PPG和加速度计数据上验证,发现符号能选择性关联健康状况和生理属性,并支持跨模态迁移。

Comments 8 pages, Mechanistic Interpretability Workshop at the 43rd International Conference on Machine Learning, 4 main figures

Journal ref Mechanistic Interpretability Workshop at the 43 rd International Conference on Machine Learning, Seoul, South Korea, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17033 2026-07-21 cs.LG physics.chem-ph 新提交 71%

ChemFusion: A Multimodal Cross-Attention Network for Reaction Yield Prediction

ChemFusion:用于反应产率预测的多模态交叉注意力网络

Qiwei Han, Chi Zhou

机构 * Duke University(杜克大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multimodal(title)

AI总结 研究过渡金属催化反应产率预测难题,提出ChemFusion多模态交叉注意力网络,融合电子特征与3D原子坐标,用交叉注意力机制,在交叉偶联库基准测试中性能出色,还能自主学习识别和惩罚空间位阻,提供物理可解释性。

Comments 10 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06305 2026-07-08 astro-ph.IM 新提交 71%

Exploring Image-Text Alignment for Radio Galaxy Morphologies

探索射电星系形态的图像-文本对齐

Erica Lastufka, Mariia Drozdova, Svyatoslav Volosynovskiy

专题命中 多模态训练与对齐 :image-text(title)

AI总结 研究射电星系图像与文本描述的对齐,利用MiraBest数据集及SigLIP-2模型,经特定提示生成描述并评估。结果显示基于描述的星系分类与图像类似,微调改善局部连贯性,此为射电星系形态研究提供新方法。

Comments Accepted at the AI in Science (AIS) Conference 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02600 2026-07-07 eess.SP 新提交 71%

A Unified Multi-Modal Sensing and Active-Stabilization Framework for Autonomous IoT Nodes in Connectivity-Denied Environments: From Perimeter Sentinel to High-Value Cargo Protection

用于连接受限环境中自主物联网节点的统一多模态传感与主动稳定框架:从周边哨兵到高价值货物保护

Naahi Mumtaj Rihan

专题命中 多模态训练与对齐 :multi-modal(title)

AI总结 探讨连接受限环境下自主物联网节点需求,提出广义节点架构,扩展其传感与通信模态,通过复合风险指数耦合遥测与地理信息系统,统一两种部署模式并给出分析与实例。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04378 2026-07-07 cs.RO 新提交 71%

SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy

SurgAM:用于机器人自主的多模态特征融合手术能力地图预测

Lei Song, Yonghao Long, Mengya Xu, Jiayi Geng, Xiuyuan Chen, Qi Dou

机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系) Department of Thoracic Surgery, Peking University People’s Hospital(北京大学人民医院胸外科) Thoracic Oncology Institute, Peking University People’s Hospital(北京大学人民医院胸部肿瘤研究所) Research Unit of Intelligence Diagnosis and Treatment in Early Non-small Cell Lung Cancer, Chinese Academy of Medical Sciences(中国医学科学院早期非小细胞肺癌智能诊断与治疗研究组) Institute of Advanced Clinical Medicine, Peking University(北京大学先进临床医学院) Beijing Key Laboratory of Innovative Application of Big Data in Lung Cancer, Peking University People’s Hospital(北京大学人民医院肺癌大数据创新应用北京市重点实验室)

专题命中 多模态训练与对齐 :multimodal(title)

AI总结 研究如何通过视觉数据识别手术可操作区域,提出自适应特征融合框架、分层提示学习机制和场景引导注意力解码器,建立新数据集验证,在真实模型上验证框架对下游自动化的适用性。

Comments ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23534 2026-06-23 eess.SP 新提交 71%

Sensor-Stack Limits on Contactless In-Bed Body Position: A 20-Subject Multimodal Radar + Thermal LOSO Characterization

传感器堆栈对非接触式床上身体位置的限制:20名受试者多模态雷达+热成像LOSO表征

Dovy Paukstys

专题命中 多模态训练与对齐 :multimodal(title)

AI总结 通过20名受试者的多模态雷达与热成像数据,研究非接触式床上体位推断的传感器表示限制,发现融合雷达与热成像的Logistic回归在床内/外分类中达到0.871中位平衡准确率,但俯卧检测性能不足(召回率0.50,精确率0.41),指出原始距离-FFT访问是下一步硬件实验方向。

Comments 11 pages, 3 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18576 2026-04-28 cs.RO 71%

DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment

DriVerse:通过多模态轨迹提示和运动对齐实现驾驶模拟的导航世界模型

Xiaofan Li, Chenming Wu, Zhao Yang, Zhihao Xu, Dingkang Liang, Yumeng Zhang, Ji Wan, Jun Wang

机构 * Baidu Inc.(百度公司)

专题命中 多模态训练与对齐 :multimodal(title)

AI总结 DriVerse通过多模态轨迹提示和运动对齐技术,实现从单张图像和未来轨迹生成导航驱动的驾驶场景,提升了动态对象的生成精度和时间一致性。

Comments 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏