arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2503.20011 2025-03-27 cs.CV cs.RO 79%

Hyperdimensional Uncertainty Quantification for Multimodal Uncertainty Fusion in Autonomous Vehicles Perception

Luke Chen, Junyao Wang, Trier Mortlock, Pramod Khargonekar, Mohammad Abdullah Al Faruque

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08497 2025-03-27 cs.LG cs.CV 79%

MMRL: Multi-Modal Representation Learning for Vision-Language Models

Yuncheng Guo, Xiaodong Gu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10636 2025-03-25 cs.LG cs.AI 79%

Adapt-$\infty$: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection

Adyasha Maharana, Jaehong Yoon, Tianlong Chen, Mohit Bansal

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments First two authors contributed equally. Code: https://github.com/adymaharana/adapt-inf

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17337 2025-03-25 cs.CV 79%

Neural-MCRL: Neural Multimodal Contrastive Representation Learning for EEG-based Visual Decoding

Yueyang Li, Zijian Kang, Shengyu Gong, Wenhao Dong, Weiming Zeng, Hongjie Yan, Wai Ting Siok, Nizhuan Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16467 2025-03-24 cs.HC cs.AI cs.RO 79%

Enhancing Explainability with Multimodal Context Representations for Smarter Robots

Anargh Viswanath, Lokesh Veeramacheneni, Hendrik Buschmeier

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Presented at 3rd Workshop on Explainability in Human-Robot Collaboration at HRI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21982 2025-03-24 cs.CV 79%

A Survey on RGB, 3D, and Multimodal Approaches for Unsupervised Industrial Image Anomaly Detection

Yuxuan Lin, Yang Chang, Xuan Tong, Jiawen Yu, Antonio Liotta, Guofan Huang, Wei Song, Deyu Zeng, Zongze Wu, Yan Wang, Wenqiang Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15022 2025-03-20 cs.CV 79%

xMOD: Cross-Modal Distillation for 2D/3D Multi-Object Discovery from 2D motion

Saad Lahlali, Sandra Kara, Hejer Ammar, Florian Chabot, Nicolas Granger, Hervé Le Borgne, Quoc-Cuong Pham

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13814 2025-03-19 cs.CV 79%

FusDreamer: Label-efficient Remote Sensing World Model for Multimodal Data Classification

Jinping Wang, Weiwei Song, Hao Chen, Jinchang Ren, Huimin Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09496 2025-03-19 cs.CV 79%

Robust Multimodal Survival Prediction with the Latent Differentiation Conditional Variational AutoEncoder

Junjie Zhou, Jiao Tang, Yingli Zuo, Peng Wan, Daoqiang Zhang, Wei Shao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11347 2025-03-18 cs.CV 79%

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

Guankun Wang, Long Bai, Junyi Wang, Kun Yuan, Zhen Li, Tianxu Jiang, Xiting He, Jinlin Wu, Zhen Chen, Zhen Lei, Hongbin Liu, Jiazheng Wang, Fan Zhang, Nicolas Padoy, Nassir Navab, Hongliang Ren

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07496 2025-03-17 cs.CV 79%

Aligning First, Then Fusing: A Novel Weakly Supervised Multimodal Violence Detection Method

Wenping Jin, Li Zhu, Jing Sun

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10057 2025-03-14 cs.CV 79%

Multi-Modal Mamba Modeling for Survival Prediction (M4Survive): Adapting Joint Foundation Model Representations

Ho Hin Lee, Alberto Santamaria-Pang, Jameson Merkov, Matthew Lungren, Ivan Tarapov

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04851 2025-03-14 cs.CV 79%

AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction

Anjun Chen, Xiangyu Wang, Zhi Xu, Kun Shi, Yan Qin, Yuchi Huo, Jiming Chen, Qi Ye

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments TMM 2025, Project Page: https://chen3110.github.io/adaptivefusion/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01131 2025-03-14 cs.CV 79%

M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression Comprehension

Xuyang Liu, Ting Liu, Siteng Huang, Yi Xin, Yue Hu, Quanjun Yin, Donglin Wang, Yuanyuan Wu, Honggang Chen

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08726 2025-03-13 cs.LG cs.AI eess.SP 79%

SIMAC: A Semantic-Driven Integrated Multimodal Sensing And Communication Framework

Yubo Peng, Luping Xiang, Kun Yang, Feibo Jiang, Kezhi Wang, Dapeng Oliver Wu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08461 2025-03-12 cs.MM cs.DC 79%

FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework

Jianian Zhu, Hang Wu, Haojie Wang, Yinghui Li, Biao Hou, Ruixuan Li, Jidong Zhai

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.MM

Comments 14 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08092 2025-03-12 cs.CV 79%

SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection

Hyeongseok Son, Jia He, Seung-In Park, Ying Min, Yunhao Zhang, ByungIn Yoo

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07655 2025-03-12 cs.LG cs.AI 79%

GraphT5: Unified Molecular Graph-Language Modeling via Multi-Modal Cross-Token Attention

Sangyeup Kim, Nayeon Kim, Yinhua Piao, Sun Kim

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07636 2025-03-12 cs.LG cs.AI 79%

An Optimization Algorithm for Multimodal Data Alignment

Wei Zhang, Xinyue Wang, Lan Yu, Shi Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments ACL SRW submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06945 2025-03-11 eess.IV cs.CV 79%

Dynamic Cross-Modal Feature Interaction Network for Hyperspectral and LiDAR Data Classification

Junyan Lin, Feng Gap, Lin Qi, Junyu Dong, Qian Du, Xinbo Gao

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE TGRS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06141 2025-03-11 cs.CV 79%

Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model

Mingxing Li, Rui Wang, Lei Sun, Yancheng Bai, Xiangxiang Chu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04444 2025-03-07 cs.CV 79%

ToFu: Visual Tokens Reduction via Fusion for Multi-modal, Multi-patch, Multi-image Task

Vittorio Pippi, Matthieu Guillaumin, Silvia Cascianelli, Rita Cucchiara, Maximilian Jaritz, Loris Bazzani

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00299 2025-03-07 cs.CV 79%

GSPR: Multimodal Place Recognition Using 3D Gaussian Splatting for Autonomous Driving

Zhangshuo Qi, Junyi Ma, Jingyi Xu, Zijie Zhou, Luqi Cheng, Guangming Xiong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03280 2025-03-06 cs.CV 79%

BEVMOSNet: Multimodal Fusion for BEV Moving Object Segmentation

Hiep Truong Cong, Ajay Kumar Sigatapu, Arindam Das, Yashwanth Sharma, Venkatesh Satagopan, Ganesh Sistu, Ciaran Eising

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments In Proceedings of the 20th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02221 2025-03-05 cs.AI 79%

Attention Bootstrapping for Multi-Modal Test-Time Adaptation

Yusheng Zhao, Junyu Luo, Xiao Luo, Jinsheng Huang, Jingyang Yuan, Zhiping Xiao, Ming Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02702 2025-03-05 cs.CV 79%

VoxelNextFusion: A Simple, Unified and Effective Voxel Fusion Framework for Multi-Modal 3D Object Detection

Ziying Song, Guoxin Zhang, Jun Xie, Lin Liu, Caiyan Jia, Shaoqing Xu, Zhepeng Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Journal ref IEEE Transactions on Geoscience and Remote Sensing, vol. 61, 2023, pp. 1-12

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01654 2025-03-04 cs.CV 79%

A Shared Encoder Approach to Multimodal Representation Learning

Shuvendu Roy, Franklin Ogidi, Ali Etemad, Elham Dolatabadi, Arash Afkanpour

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01022 2025-03-04 cond-mat.mtrl-sci cs.AI cs.LG 79%

LLM-Fusion: A Novel Multimodal Fusion Model for Accelerated Material Discovery

Onur Boyar, Indra Priyadarsini, Seiji Takeda, Lisa Hamada

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 4 pages, presented at AAAI 2025 Workshop on AI to Accelerating Science and Engineering (AI2ASE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00364 2025-03-04 cs.CV 79%

CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion

Yaowei Guo, Jiazheng Xing, Xiaojun Hou, Shuo Xin, Juntao Jiang, Demetri Terzopoulos, Chenfanfu Jiang, Yong Liu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20077 2025-03-03 cs.CV 79%

SegLocNet: Multimodal Localization Network for Autonomous Driving via Bird's-Eye-View Segmentation

Zijie Zhou, Zhangshuo Qi, Luqi Cheng, Guangming Xiong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏