arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2304.08709 2024-03-25 cs.CV 79%

You Only Need Two Detectors to Achieve Multi-Modal 3D Multi-Object Tracking

Xiyang Wang, Chunyun Fu, Jiawei He, Mingguang Huang, Ting Meng, Siyu Zhang, Hangning Zhou, Ziyao Xu, Chi Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13830 2024-03-22 q-bio.BM cs.CL cs.LG 79%

Bridging Text and Molecule: A Survey on Multimodal Frameworks for Molecule

Yi Xiao, Xiangxin Zhou, Qiang Liu, Liang Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12725 2024-03-21 cs.CL 79%

Generative Multimodal Entity Linking

Senbao Shi, Zhenran Xu, Baotian Hu, Min Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10036 2024-03-18 cs.CV 79%

SparseFusion: Efficient Sparse Multi-Modal Fusion Framework for Long-Range 3D Perception

Yiheng Li, Hongyang Li, Zehao Huang, Hong Chang, Naiyan Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08182 2024-03-14 cs.CV 79%

SeCG: Semantic-Enhanced 3D Visual Grounding via Cross-modal Graph Attention

Feng Xiao, Hongbin Xu, Qiuxia Wu, Wenxiong Kang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04829 2024-03-14 cs.CV 79%

MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation

Kaixin Cai, Pengzhen Ren, Yi Zhu, Hang Xu, Jianzhuang Liu, Changlin Li, Guangrun Wang, Xiaodan Liang

专题命中 多模态训练与对齐 :cross-modal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07301 2024-03-13 cs.CV 79%

Let Storytelling Tell Vivid Stories: An Expressive and Fluent Multimodal Storyteller

Chuanqi Zang, Jiji Tang, Rongsheng Zhang, Zeng Zhao, Tangjie Lv, Mingtao Pei, Wei Liang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10601 2024-03-13 cs.CV eess.SP 79%

Multimodal Indoor Localization Using Crowdsourced Radio Maps

Zhaoguang Yi, Xiangyu Wen, Qiyue Xia, Peize Li, Francisco Zampella, Firas Alsehly, Chris Xiaoxuan Lu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 4 figures; ICRA'24 https://youtu.be/NTTKwJBFN5w

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05552 2024-03-12 cs.CY cs.AI cs.LG 79%

Multi-source and multimodal data fusion for predicting academic performance in blended learning university courses

W. Chango, R. Cerezo, C. Romero

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Journal ref Computers & Electrical Engineering, 89, 106908 (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06339 2024-03-12 cs.CV 79%

FOAA: Flattened Outer Arithmetic Attention For Multimodal Tumor Classification

Omnia Alwazzan, Ioannis Patras, Gregory Slabaugh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments This paper has been accepted for ISBI-2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01802 2024-03-12 cs.CV 79%

TNF: Tri-branch Neural Fusion for Multimodal Medical Data Classification

Tong Zheng, Shusaku Sone, Yoshitaka Ushiku, Yuki Oba, Jiaxin Ma

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03217 2024-03-06 cs.CV 79%

Self-supervised 3D Patient Modeling with Multi-modal Attentive Fusion

Meng Zheng, Benjamin Planche, Xuan Gong, Fan Yang, Terrence Chen, Ziyan Wu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments MICCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02991 2024-03-06 cs.CV 79%

MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer

Jianjian Cao, Peng Ye, Shengze Li, Chong Yu, Yansong Tang, Jiwen Lu, Tao Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 19 pages, 9 figures, Published in CVPR2024

Journal ref In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01203 2024-03-05 cs.LG cs.CL cs.DB 79%

Pseudo-Label Calibration Semi-supervised Multi-Modal Entity Alignment

Luyao Wang, Pengnian Qi, Xigang Bao, Chunlai Zhou, Biao Qin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

Comments accepted by AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17936 2024-02-29 cs.CL 79%

Acquiring Linguistic Knowledge from Multimodal Input

Theodor Amariucai, Alex Warstadt

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments in Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09831 2024-02-29 eess.IV cs.CV 79%

Cross-modality Attention-based Multimodal Fusion for Non-small Cell Lung Cancer (NSCLC) Patient Survival Prediction

Ruining Deng, Nazim Shaikh, Gareth Shannon, Yao Nie

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17483 2024-02-28 cs.CV 79%

AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis

Tao Tang, Guangrun Wang, Yixing Lao, Peng Chen, Jie Liu, Liang Lin, Kaicheng Yu, Xiaodan Liang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11826 2024-02-20 cs.CV 79%

Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios

Jialei Xu, Xianming Liu, Junjun Jiang, Kui Jiang, Rui Li, Kai Cheng, Xiangyang Ji

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11307 2024-02-20 cs.CV 79%

ICHPro: Intracerebral Hemorrhage Prognosis Classification Via Joint-attention Fusion-based 3d Cross-modal Network

Xinlei Yu, Xinyang Li, Ruiquan Ge, Shibin Wu, Ahmed Elazab, Jichao Zhu, Lingyan Zhang, Gangyong Jia, Taosheng Xu, Xiang Wan, Changmiao Wang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments 6 pages,4 figures, 4 tables, accepted by ISBI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07358 2024-02-19 cs.CL 79%

Towards Versatile and Efficient Visual Knowledge Integration into Pre-trained Language Models with Cross-Modal Adapters

Xinyun Zhang, Haochen Tan, Han Wu, Bei Yu

专题命中 多模态训练与对齐 :cross-modal(title);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09738 2024-02-16 cs.CL 79%

Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection

Eftekhar Hossain, Omar Sharif, Mohammed Moshiul Hoque, Sarah M. Preum

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to EACL-SRW, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02384 2024-02-16 cs.CV 79%

ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning

Fanqing Meng, Wenqi Shao, Quanfeng Lu, Peng Gao, Kaipeng Zhang, Yu Qiao, Ping Luo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Updated and corrected experimental results, removal of inappropriate experiments, and a more comprehensive experimental setup

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05151 2024-02-09 cs.LG cs.AI 79%

CrashFormer: A Multimodal Architecture to Predict the Risk of Crash

Amin Karimi Monsefi, Pouya Shiri, Ahmad Mohammadshirazi, Nastaran Karimi Monsefi, Ron Davies, Sobhan Moosavi, Rajiv Ramnath

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.AI

Comments The paper is accepted In 1st ACM SIGSPATIAL International Workshop on Advances in Urban-AI (UrbanAI 23), November 13, 2023, Hamburg, Germany

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02055 2024-02-06 cs.LG cs.AI 79%

Variance Alignment Score: A Simple But Tough-to-Beat Data Selection Method for Multimodal Contrastive Learning

Yiping Wang, Yifang Chen, Wendan Yan, Kevin Jamieson, Simon Shaolei Du

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.AI

Comments 17 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13508 2024-02-06 cs.LG cs.AI cs.DC 79%

Multimodal Federated Learning with Missing Modality via Prototype Mask and Contrast

Guangyin Bao, Qi Zhang, Duoqian Miao, Zixuan Gong, Liang Hu, Ke Liu, Yang Liu, Chongyang Shi

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01311 2024-02-05 cs.CV eess.IV 79%

Deep Multimodal Fusion of Data with Heterogeneous Dimensionality via Projective Networks

José Morano, Guilherme Aresta, Christoph Grechenig, Ursula Schmidt-Erfurth, Hrvoje Bogunović

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted for publication in the IEEE Journal of Biomedical and Health Informatics (JBHI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01886 2024-02-01 cs.CV 79%

Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image Fusion

Xilai Li, Xiaosong Li, Tao Ye, Xiaoqi Cheng, Wuyang Liu, Haishu Tan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09479 2024-01-24 cs.CR cs.AI cs.LG 79%

Uncertainty-Aware Hardware Trojan Detection Using Multimodal Deep Learning

Rahul Vishwakarma, Amin Rezaei

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 2024 Design, Automation and Test in Europe Conference | The European Event for Electronic System Design & Test (accepted)

Journal ref 2024 Design, Automation and Test in Europe Conference | The European Event for Electronic System Design & Test

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03179 2024-01-24 cs.CV 79%

Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR Classification

Jiaqing Zhang, Jie Lei, Weiying Xie, Geng Yang, Daixun Li, Yunsong Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11740 2024-01-23 cs.CV cs.LG 79%

Multi-level Cross-modal Alignment for Image Clustering

Liping Qiu, Qin Zhang, Xiaojun Chen, Shaotian Cai

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏