arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2411.18428 2025-01-03 cs.LG cs.AI 83%

MM-Path: Multi-modal, Multi-granularity Path Representation Learning -- Extended Version

Ronghui Xu, Hanyin Cheng, Chenjuan Guo, Hongfan Gao, Jilin Hu, Sean Bin Yang, Bin Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments This is an extended version of the paper accepted by KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18327 2025-01-03 eess.IV cs.CV cs.LG 83%

Multi-modal Evidential Fusion Network for Trustworthy PET/CT Tumor Segmentation

Yuxuan Qi, Li Lin, Jiajun Wang, Bin Zhang, Jingya Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18176 2024-12-31 cs.IR cs.AI 83%

Molar: Multimodal LLMs with Collaborative Filtering Alignment for Enhanced Sequential Recommendation

Yucong Luo, Qitao Qin, Hao Zhang, Mingyue Cheng, Ruiran Yan, Kefan Wang, Jie Ouyang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08827 2024-12-31 cs.CV 83%

RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba

Andong Lu, Wanyu Wang, Chenglong Li, Jin Tang, Bin Luo

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19128 2024-12-30 cs.CV cs.LG 83%

Semantic Residual for Multimodal Unified Discrete Representation

Hai Huang, Shulei Wang, Yan Xia

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICASSP 2025 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13637 2024-12-30 cs.CV 83%

Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation

Sen Lei, Xinyu Xiao, Tianlin Zhang, Heng-Chao Li, Zhenwei Shi, Qing Zhu

专题命中 多模态训练与对齐 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by IEEE TGRS

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18124 2024-12-25 cs.CV 83%

VisionLLM-based Multimodal Fusion Network for Glottic Carcinoma Early Detection

Zhaohui Jin, Yi Shuai, Yongcheng Li, Lingcong Cai, Yun Li, Huifen Liu, Xiaomao Fan

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19101 2024-12-20 cs.CV 83%

DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Jiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie, Lianwen Jin

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09468 2024-12-17 cs.AI 83%

Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity Representation

Yichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu, Binbin Hu, Ziqi Liu, Wen Zhang, Huajun Chen

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments AAAI 2025; Repo is available at https://github.com/zjukg/MyGO

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09870 2024-12-16 cs.CV 83%

Dynamic Cross-Modal Alignment for Robust Semantic Location Prediction

Liu Jing, Amirul Rahman

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19896 2024-12-10 cs.CV 83%

FLAASH: Flow-Attention Adaptive Semantic Hierarchical Fusion for Multi-Modal Tobacco Content Analysis

Naga VS Raviteja Chappa, Page Daniel Dobbs, Bhiksha Raj, Khoa Luu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Under review at Image and Vision Computing Journal; 20 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05660 2024-12-10 cs.CV 83%

Multimodal Biometric Authentication Using Camera-Based PPG and Fingerprint Fusion

Xue Xian Zheng, M. M. Ur Rahma, Bilal Taha, Mudassir Masood, Dimitrios Hatzinakos, Tareq Al-Naffouri

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00556 2024-12-10 cs.CV 83%

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction

Shiyu Zhao, Zhenting Wang, Felix Juefei-Xu, Xide Xia, Miao Liu, Xiaofang Wang, Mingfu Liang, Ning Zhang, Dimitris N. Metaxas, Licheng Yu

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Technical report, 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03625 2024-12-06 cs.CL 83%

Multimodal Sentiment Analysis Based on BERT and ResNet

JiaLe Ren

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02283 2024-12-04 eess.SP cs.AI 83%

VR Based Emotion Recognition Using Deep Multimodal Fusion With Biosignals Across Multiple Anatomical Domains

Pubudu L. Indrasiri, Bipasha Kashyap, Chandima Kolambahewage, Bahareh Nakisa, Kiran Ijaz, Pubudu N. Pathirana

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06169 2024-12-03 cs.CV 83%

Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Zeliang Zhang, Phu Pham, Wentian Zhao, Kun Wan, Yu-Jhe Li, Jianing Zhou, Daniel Miranda, Ajinkya Kale, Chenliang Xu

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13909 2024-11-25 cs.CV 83%

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts

Honglin Li, Yuting Gao, Chenglu Zhu, Jingdong Chen, Ming Yang, Lin Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04697 2024-11-08 cs.CV 83%

Dynamic Brightness Adaptation for Robust Multi-modal Image Fusion

Yiming Sun, Bing Cao, Pengfei Zhu, Qinghua Hu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by IJCAI 2024

Journal ref Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence,Main Track,Pages 1317-1325, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16591 2024-11-08 cs.CV 83%

CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification

Qijie Wang, Guandu Liu, Bin Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ACM Multimedia 2024 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23089 2024-10-31 cs.CV 83%

PIP-MM: Pre-Integrating Prompt Information into Visual Encoding via Existing MLLM Structures

Tianxiang Wu, Minxin Nie, Ziqiang Cao

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09115 2024-10-29 cs.LG cs.AI 83%

HEALNet: Multimodal Fusion for Heterogeneous Biomedical Data

Konstantin Hemker, Nikola Simidjievski, Mateja Jamnik

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19446 2024-10-28 cs.CV 83%

Fusion-then-Distillation: Toward Cross-modal Positive Distillation for Domain Adaptive 3D Semantic Segmentation

Yao Wu, Mingwei Xing, Yachao Zhang, Yuan Xie, Yanyun Qu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11402 2024-10-24 cs.CL cs.AI cs.CV cs.LG cs.MM 83%

NVLM: Open Frontier-Class Multimodal LLMs

Wenliang Dai, Nayeon Lee, Boxin Wang, Zhuolin Yang, Zihan Liu, Jon Barker, Tuomas Rintamaki, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Fixed the typos. For more information, please visit our project page at: https://research.nvidia.com/labs/adlr/NVLM-1

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16853 2024-10-23 cs.CV cs.IR 83%

Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text Matching

Xiang Ma, Xuemei Li, Lexin Fang, Caiming Zhang

专题命中 多模态训练与对齐 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15517 2024-10-22 cs.CL 83%

SceneGraMMi: Scene Graph-boosted Hybrid-fusion for Multi-Modal Misinformation Veracity Prediction

Swarang Joshi, Siddharth Mavani, Joel Alex, Arnav Negi, Rahul Mishra, Ponnurangam Kumaraguru

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15015 2024-10-22 cs.CV 83%

MambaSOD: Dual Mamba-Driven Cross-Modal Fusion Network for RGB-D Salient Object Detection

Yue Zhan, Zhihong Zeng, Haijun Liu, Xiaoheng Tan, Yinli Tian

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14944 2024-10-22 cs.CV 83%

Part-Whole Relational Fusion Towards Multi-Modal Scene Understanding

Yi Liu, Chengxin Li, Shoukun Xu, Jungong Han

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21659 2024-10-18 cs.CL 83%

Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models

Yue Xu, Xiuyuan Qi, Zhan Qin, Wenjie Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments 12 pages, 9 figures, EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11358 2024-10-16 cs.CV 83%

SeaDATE: Remedy Dual-Attention Transformer with Semantic Alignment via Contrast Learning for Multimodal Object Detection

Shuhan Dong, Yunsong Li, Weiying Xie, Jiaqing Zhang, Jiayuan Tian, Danian Yang, Jie Lei

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09572 2024-10-16 cs.CV 83%

Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation

Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung, James T. Kwok, Yu Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ECCV2024 (Project Page: https://gyhdog99.github.io/projects/ecso/)

详情

展开后加载摘要…

URL PDF HTML 收藏