arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2405.07702 2024-05-14 cs.CV cs.LG 79%

FORESEE: Multimodal and Multi-view Representation Learning for Robust Prediction of Cancer Survival

Liangrui Pan, Yijun Peng, Yan Li, Yiyi Liang, Liwen Xu, Qingchun Liang, Shaoliang Peng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.06128 2024-05-13 cs.CV 79%

Enhanced Multimodal Content Moderation of Children's Videos using Audiovisual Fusion

Syed Hammad Ahmed, Muhammad Junaid Khan, Gita Sukthankar

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 3 figures, Accepted at The 37th International FLAIRS Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14516 2024-05-09 cs.CV cs.RO 79%

UniBEV: Multi-modal 3D Object Detection with Uniform BEV Encoders for Robustness against Missing Sensor Modalities

Shiming Wang, Holger Caesar, Liangliang Nan, Julian F. P. Kooij

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Intelligent Vehicles Symposium (IV 2024), camera-ready. Code: https://github.com/tudelft-iv/UniBEV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02649 2024-05-07 cs.LG cs.AI 79%

Generic Multi-modal Representation Learning for Network Traffic Analysis

Luca Gioacchini, Idilio Drago, Marco Mellia, Zied Ben Houidi, Dario Rossi

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17858 2024-05-06 cs.CL 79%

Revisiting Multi-modal Emotion Learning with Broad State Space Models and Probability-guidance Fusion

Yuntao Shou, Tao Meng, Fuchen Zhang, Nan Yin, Keqin Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16233 2024-05-02 cs.LG cs.AI 79%

AutoGluon-Multimodal (AutoMM): Supercharging Multimodal AutoML with Foundation Models

Zhiqiang Tang, Haoyang Fang, Su Zhou, Taojiannan Yang, Zihan Zhong, Tony Hu, Katrin Kirchhoff, George Karypis

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at AutoML 2024 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11210 2024-05-02 cs.CV 79%

HiH: A Multi-modal Hierarchy in Hierarchy Network for Unconstrained Gait Recognition

Lei Wang, Bo Liu, Yinchi Ma, Fangfang Liang, Nawei Guo

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17590 2024-04-30 cs.IR cs.AI 79%

Leveraging Intra-modal and Inter-modal Interaction for Multi-Modal Entity Alignment

Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li, Jeff Z. Pan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08275 2024-04-29 cs.CV 79%

ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding

Le Xue, Ning Yu, Shu Zhang, Artemis Panagopoulou, Junnan Li, Roberto Martín-Martín, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, Silvio Savarese

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments CVPR2024

Journal ref CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15877 2024-04-24 cs.LG cs.AI 79%

Neuro-Inspired Hierarchical Multimodal Learning

Xiongye Xiao, Gengshuo Liu, Gaurav Gupta, Defu Cao, Shixuan Li, Yaxing Li, Tianqing Fang, Mingxi Cheng, Paul Bogdan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments I am requesting the withdrawal of this submission due to an inadvertent duplication. The paper was submitted twice under different IDs, which was not intentional. The other submission (arXiv:2404.09403) contains the most updated and comprehensive version of the paper, and I would like to retain that as the sole version on the platform

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13596 2024-04-24 eess.IV eess.AS eess.SP 79%

Multimodal sensor fusion for real-time location-dependent defect detection in laser-directed energy deposition

Lequn Chen, Xiling Yao, Wenhe Feng, Youxiang Chew, Seung Ki Moon

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 eess.AS

Comments 8 pages, 10 figures. This paper has been accepted to be published in the proceedings of IDETC-CIE 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13704 2024-04-23 eess.IV cs.CV cs.LG 79%

PEMMA: Parameter-Efficient Multi-Modal Adaptation for Medical Image Segmentation

Nada Saadi, Numan Saeed, Mohammad Yaqub, Karthik Nandakumar

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13619 2024-04-23 cs.MM 79%

Towards Unified Representation of Multi-Modal Pre-training for 3D Understanding via Differentiable Rendering

Ben Fei, Yixuan Li, Weidong Yang, Lipeng Ma, Ying He

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12588 2024-04-22 cs.CV cs.LG 79%

Cross-Modal Adapter: Parameter-Efficient Transfer Learning Approach for Vision-Language Models

Juncheng Yang, Zuchao Li, Shuai Xie, Weiping Zhu, Wei Yu, Shijun Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments This paper is accepted to ICME 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16666 2024-04-22 cs.LG cs.AI physics.chem-ph q-bio.BM 79%

MultiModal-Learning for Predicting Molecular Properties: A Framework Based on Image and Graph Structures

Zhuoyuan Wang, Jiacong Mi, Shan Lu, Jieyue He

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05927 2024-04-22 cs.LG cs.AI eess.SP 79%

Frequency-Aware Masked Autoencoders for Multimodal Pretraining on Biosignals

Ran Liu, Ellen L. Zippi, Hadi Pouransari, Chris Sandino, Jingping Nie, Hanlin Goh, Erdrin Azemi, Ali Moin

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Extended version of ICLR 2024 Learning from Time Series for Health workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11764 2024-04-19 cs.CV 79%

Multimodal 3D Object Detection on Unseen Domains

Deepti Hegde, Suhas Lohit, Kuan-Chuan Peng, Michael J. Jones, Vishal M. Patel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10381 2024-04-18 cs.IR cs.AI 79%

UMAIR-FPS: User-aware Multi-modal Animation Illustration Recommendation Fusion with Painting Style

Yan Kang, Hao Lin, Mingjian Yang, Shin-Jye Lee

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by DASFAA 2024 Research track

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10146 2024-04-17 cs.CV 79%

Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels

Amaya Dharmasiri, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments To be published in Workshop for Learning 3D with Multi-View Supervision (3DMV) at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09499 2024-04-16 cs.CV cs.GR 79%

Learning Human Motion from Monocular Videos via Cross-Modal Manifold Alignment

Shuaiying Hou, Hongyu Tao, Junheng Fang, Changqing Zou, Hujun Bao, Weiwei Xu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08923 2024-04-16 cs.CV 79%

Trustworthy Multimodal Fusion for Sentiment Analysis in Ordinal Sentiment Space

Zhuyang Xie, Yan Yang, Jie Wang, Xiaorong Liu, Xiaofan Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 14 pages, 9 figures, Accepted by IEEE Transactions on Circuits and Systems for Video Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06791 2024-04-16 cs.CV 79%

PV-SSD: A Multi-Modal Point Cloud Feature Fusion Method for Projection Features and Variable Receptive Field Voxel Features

Yongxin Shao, Aihong Tan, Zhetao Sun, Enhui Zheng, Tianhong Yan, Peng Liao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06107 2024-04-10 cs.CL 79%

Exploring the Necessity of Visual Modality in Multimodal Machine Translation using Authentic Datasets

Zi Long, Zhenhao Tang, Xianghua Fu, Jian Chen, Shilong Hou, Jinze Lyu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments bucc 2024 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16119 2024-04-09 cs.AI 79%

Triple Disentangled Representation Learning for Multimodal Affective Analysis

Ying Zhou, Xuefeng Liang, Han Chen, Yin Zhao, Xin Chen, Lida Yu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04026 2024-04-08 cs.RO cs.CV 79%

MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes

Chenyang Wu, Yifan Duan, Xinran Zhang, Yu Sheng, Jianmin Ji, Yanyong Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00144 2024-04-02 eess.IV cs.CV 79%

An Interpretable Cross-Attentive Multi-modal MRI Fusion Framework for Schizophrenia Diagnosis

Ziyu Zhou, Anton Orlichenko, Gang Qu, Zening Fu, Vince D Calhoun, Zhengming Ding, Yu-Ping Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16002 2024-03-29 cs.CV 79%

SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

Xiaojun Hou, Jiazheng Xing, Yijie Qian, Yaowei Guo, Shuo Xin, Junhao Chen, Kai Tang, Mengmeng Wang, Zhengkai Jiang, Liang Liu, Yong Liu

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16257 2024-03-26 cs.CV 79%

Unlearning Backdoor Threats: Enhancing Backdoor Defense in Multimodal Contrastive Learning via Local Token Unlearning

Siyuan Liang, Kuanrong Liu, Jiajun Gong, Jiawei Liang, Yuan Xun, Ee-Chien Chang, Xiaochun Cao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01099 2024-03-26 eess.IV cs.CV cs.LG 79%

HyMNet: a Multimodal Deep Learning System for Hypertension Classification using Fundus Photographs and Cardiometabolic Risk Factors

Mohammed Baharoon, Hessa Almatar, Reema Alduhayan, Tariq Aldebasi, Badr Alahmadi, Yahya Bokhari, Mohammed Alawad, Ahmed Almazroa, Abdulrhman Aljouie

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15241 2024-03-25 cs.CV 79%

IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection

Junbo Yin, Jianbing Shen, Runnan Chen, Wei Li, Ruigang Yang, Pascal Frossard, Wenguan Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2024; Code: https://github.com/yinjunbo/IS-Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏