arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2406.14455 2025-03-31 cs.CV 83%

MM-GTUNets: Unified Multi-Modal Graph Deep Learning for Brain Disorders Prediction

Luhui Cai, Weiming Zeng, Hongyu Chen, Hua Zhang, Yueyang Li, Yu Feng, Hongjie Yan, Lingbin Bian, Wai Ting Siok, Nizhuan Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18100 2025-03-25 cs.CV 83%

M3Net: Multimodal Multi-task Learning for 3D Detection, Segmentation, and Occupancy Prediction in Autonomous Driving

Xuesong Chen, Shaoshuai Shi, Tao Ma, Jingqiu Zhou, Simon See, Ka Chun Cheung, Hongsheng Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15284 2025-03-20 cs.CV 83%

EdgeRegNet: Edge Feature-based Multimodal Registration Network between Images and LiDAR Point Clouds

Yuanchao Yue, Hui Yuan, Qinglong Miao, Xiaolong Mao, Raouf Hamzaoui, Peter Eisert

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08992 2025-03-18 cs.CV 83%

Dual-Domain Homogeneous Fusion with Cross-Modal Mamba and Progressive Decoder for 3D Object Detection

Xuzhong Hu, Zaipeng Duan, Pei An, Jun zhang, Jie Ma

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 13 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12560 2025-03-18 cs.CL 83%

Multi-Granular Multimodal Clue Fusion for Meme Understanding

Li Zheng, Hao Fei, Ting Dai, Zuquan Peng, Fei Li, Huisheng Ma, Chong Teng, Donghong Ji

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01217 2025-03-18 cs.CV 83%

CromSS: Cross-modal pre-training with noisy labels for remote sensing image segmentation

Chenying Liu, Conrad Albrecht, Yi Wang, Xiao Xiang Zhu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments The 1st short version was accepted as an oral presentation by ICLR 2024 ML4RS workshop. The 2nd extended version was accepted by IEEE TGRS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11183 2025-03-17 cs.CV 83%

Multimodal-Aware Fusion Network for Referring Remote Sensing Image Segmentation

Leideng Shi, Juan Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 5 pages, 5 figures, accepted in IEEE Geoscience and Remote Sensing Letters (GRSL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10501 2025-03-14 cs.CV 83%

TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models

Xudong Tan, Peng Ye, Chongjun Tu, Jianjian Cao, Yaoxin Yang, Lin Zhang, Dongzhan Zhou, Tao Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09873 2025-03-14 cs.CV 83%

FDCT: Frequency-Aware Decomposition and Cross-Modal Token-Alignment for Multi-Sensor Target Classification

Shoaib Meraj Sami, Md Mahedi Hasan, Nasser M. Nasrabadi, Raghuveer Rao

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 12 pages Accepted in the IEEE Transactions on Aerospace and Electronic Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07593 2025-03-11 cs.CV 83%

Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Youjun Zhao, Jiaying Lin, Rynson W. H. Lau

专题命中 多模态训练与对齐 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments AAAI 2025 (Extented Version). Project Page: https://youjunzhao.github.io/HCMA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13652 2025-03-11 cs.CL 83%

LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models

Yizheng Sun, Yanze Xin, Hao Li, Jingyuan Sun, Chenghua Lin, Riza Batista-Navarro

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CL

Comments Accepted to NAACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06063 2025-03-11 cs.CV 83%

Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices

Junyan Lin, Haoran Chen, Yue Fan, Yingqi Fan, Xin Jin, Hui Su, Jinlan Fu, Xiaoyu Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03202 2025-03-06 cs.CV 83%

Variance-Aware Loss Scheduling for Multimodal Alignment in Low-Data Settings

Sneh Pillai

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00591 2025-03-04 cs.CV 83%

AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models

Sohan Patnaik, Rishabh Jain, Balaji Krishnamurthy, Mausoom Sarkar

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted for publication in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12164 2025-03-04 cs.LG cs.AI 83%

GAMED: Knowledge Adaptive Multi-Experts Decoupling for Multimodal Fake News Detection

Lingzhi Shen, Yunfei Long, Xiaohao Cai, Imran Razzak, Guanming Chen, Kang Liu, Shoaib Jameel

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16786 2025-03-03 cs.CV 83%

SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding

Liangtao Shi, Ting Liu, Xiantao Hu, Yue Hu, Quanjun Yin, Richang Hong

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20008 2025-02-28 cs.CV 83%

Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up

Lang Huang, Qiyu Wu, Zhongtao Miao, Toshihiko Yamasaki

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18481 2025-02-27 cs.IR cs.AI 83%

MDE: Modality Discrimination Enhancement for Multi-modal Recommendation

Hang Zhou, Yucheng Wang, Huijing Zhan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17425 2025-02-25 cs.CV cs.LG 83%

Introducing Visual Perception Token into Multimodal Large Language Model

Runpeng Yu, Xinyin Ma, Xinchao Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13624 2025-02-20 cs.CV 83%

CardiacMamba: A Multimodal RGB-RF Fusion Framework with State Space Models for Remote Physiological Measurement

Zheng Wu, Yiping Xie, Bo Zhao, Jiguang He, Fei Luo, Ning Deng, Zitong Yu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10419 2025-02-19 cs.NE cs.AI cs.LG 83%

A Hybrid Swarm Intelligence Approach for Optimizing Multimodal Large Language Models Deployment in Edge-Cloud-based Federated Learning Environments

Gaith Rjouba, Hanae Elmekki, Saidul Islam, Jamal Bentahar, Rachida Dssouli

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10167 2025-02-18 cs.CV eess.SP 83%

X-Fi: A Modality-Invariant Foundation Model for Multimodal Human Sensing

Xinyan Chen, Jianfei Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01943 2025-02-12 cs.CV 83%

DAMA: Data- and Model-aware Alignment of Multi-modal LLMs

Jinda Lu, Junkang Wu, Jinghan Li, Xiaojun Jia, Shuo Wang, YiFan Zhang, Junfeng Fang, Xiang Wang, Xiangnan He

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06194 2025-02-11 cs.CV 83%

Multimodal Task Representation Memory Bank vs. Catastrophic Forgetting in Anomaly Detection

You Zhou, Jiangshan Zhao, Deyu Zeng, Zuo Zuo, Weixiang Liu, Zongze Wu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03499 2025-02-07 q-bio.GN cs.AI cs.LG 83%

Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning

Zehui Li, Vallijah Subasri, Yifei Shen, Dongsheng Li, Yiren Zhao, Guy-Bart Stan, Caihua Shan

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02438 2025-02-05 cs.CR cs.AI 83%

Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment

Yaling Shen, Zhixiong Zhuang, Kun Yuan, Maria-Irina Nicolae, Nassir Navab, Nicolas Padoy, Mario Fritz

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Accepted at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09498 2025-02-04 cs.CV 83%

FusionMamba: Dynamic Feature Enhancement for Multimodal Image Fusion with Mamba

Xinyu Xie, Yawen Cui, Tao Tan, Xubin Zheng, Zitong Yu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by Visual Intelligence. Codes are at https://github.com/millieXie/FusionMamba

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16621 2025-01-29 cs.LG cs.AI 83%

Chinese Stock Prediction Based on a Multi-Modal Transformer Framework: Macro-Micro Information Fusion

Lumen AI, Tengzhou No. 1 Middle School, Shihao Ji, Zihui Song, Fucheng Zhong, Jisen Jia, Zhaobo Wu, Zheyi Cao, Xu Tianhao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16998 2025-01-09 eess.IV cs.CV 83%

Deep Unfolding Network with Spatial Alignment for multi-modal MRI reconstruction

Hao Zhang, Qi Wang, Jun Shi, Shihui Ying, Zhijie Wen

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01240 2025-01-03 cs.CV 83%

Asymmetric Reinforcing against Multi-modal Representation Bias

Xiyuan Gao, Bing Cao, Pengfei Zhu, Nannan Wang, Qinghua Hu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏