arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2411.01445 2024-11-05 cs.CV 79%

A Visual Question Answering Method for SAR Ship: Breaking the Requirement for Multimodal Dataset Construction and Model Fine-Tuning

Fei Wang, Chengcheng Chen, Hongyu Chen, Yugang Chang, Weiming Zeng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01399 2024-11-05 cs.CV 79%

MambaReg: Mamba-Based Disentangled Convolutional Sparse Coding for Unsupervised Deformable Multi-Modal Image Registration

Kaiang Wen, Bin Xie, Bin Duan, Yan Yan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02905 2024-11-05 cs.LG cs.AI eess.SP 79%

H2G2-Net: A Hierarchical Heterogeneous Graph Generative Network Framework for Discovery of Multi-Modal Physiological Responses

Haidong Gu, Nathan Gaw, Yinan Wang, Chancellor Johnstone, Christine Beauchene, Sophia Yuditskaya, Hrishikesh Rao, Chun-An Chou

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments Paper accepted in Human-Centric Representation Learning workshop at AAAI 2024 (https://hcrl-workshop.github.io/2024/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00749 2024-11-04 eess.IV cs.CV q-bio.GN q-bio.TO 79%

PathoGen-X: A Cross-Modal Genomic Feature Trans-Align Network for Enhanced Survival Prediction from Histopathology Images

Akhila Krishna, Nikhil Cherian Kurian, Abhijeet Patil, Amruta Parulekar, Amit Sethi

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18947 2024-11-04 cs.LG cs.AI 79%

Multimodal Fusion on Low-quality Data: A Comprehensive Survey

Qingyang Zhang, Yake Wei, Zongbo Han, Huazhu Fu, Xi Peng, Cheng Deng, Qinghua Hu, Cai Xu, Jie Wen, Di Hu, Changqing Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Feel free to comment on our manuscript: qingyangzhang@tju$.$edu$.$cn

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23988 2024-11-01 cs.CV 79%

JEMA: A Joint Embedding Framework for Scalable Co-Learning with Multimodal Alignment

Joao Sousa, Roya Darabi, Armando Sousa, Frank Brueckner, Luís Paulo Reis, Ana Reis

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 26 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21831 2024-10-30 cs.CV 79%

Enhanced Survival Prediction in Head and Neck Cancer Using Convolutional Block Attention and Multimodal Data Fusion

Aiman Farooq, Utkarsh Sharma, Deepak Mishra

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Accepted to [ACCV 2024 Workshop]

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15659 2024-10-30 cs.CV eess.IV 79%

DeepLight: Reconstructing High-Resolution Observations of Nighttime Light With Multi-Modal Remote Sensing Data

Lixian Zhang, Runmin Dong, Shuai Yuan, Jinxiao Zhang, Mengxuan Chen, Juepeng Zheng, Haohuan Fu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This paper has been accepted in IJCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20969 2024-10-29 cs.RO cs.CV 79%

BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment

Mehdi Hosseinzadeh, Ian Reid

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted for presentation at the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024. Project page: https://m80hz.github.io/bevpose/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19974 2024-10-29 cs.LG cs.CL cs.IR 79%

Evaluating Cost-Accuracy Trade-offs in Multimodal Search Relevance Judgements

Silvia Terragni, Hoang Cuong, Joachim Daiber, Pallavi Gudipati, Pablo N. Mendes

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Journal ref CIKM MMSR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19788 2024-10-29 eess.SP cs.CV cs.LG 79%

Multi-modal Image and Radio Frequency Fusion for Optimizing Vehicle Positioning

Ouwen Huan, Tao Luo, Mingzhe Chen

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10156 2024-10-28 eess.IV cs.CV 79%

Cross-Sim-NGF: FFT-Based Global Rigid Multimodal Alignment of Image Volumes using Normalized Gradient Fields

Johan Öfverstedt, Joakim Lindblad, Nataša Sladoje

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 5 pages, 3 figures, 3 tables. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16818 2024-10-28 cs.RO cs.CV 79%

MEM: Multi-Modal Elevation Mapping for Robotics and Learning

Gian Erni, Jonas Frey, Takahiro Miki, Matias Mattamala, Marco Hutter

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accapted for IROS2023. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07981 2024-10-25 cs.LG cs.AI 79%

MolMix: A Simple Yet Effective Baseline for Multimodal Molecular Representation Learning

Andrei Manolache, Dragos Tantaru, Mathias Niepert

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Machine Learning for Structural Biology Workshop, NeurIPS 2024 v2: Added optimizer references

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03489 2024-10-24 cs.CR cs.AI 79%

Gradient-based Jailbreak Images for Multimodal Fusion Models

Javier Rando, Hannah Korevaar, Erik Brinkman, Ivan Evtimov, Florian Tramèr

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17903 2024-10-24 cs.CV q-bio.NC 79%

Reliable Object Tracking by Multimodal Hybrid Feature Extraction and Transformer-Based Fusion

Hongze Sun, Rui Liu, Wuque Cai, Jun Wang, Yue Wang, Huajin Tang, Yan Cui, Dezhong Yao, Daqing Guo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 16 pages, 7 figures, 9 tabes; This work has been submitted for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06197 2024-10-24 eess.IV cs.CV cs.LG 79%

DrFuse: Learning Disentangled Representation for Clinical Multi-Modal Fusion with Missing Modality and Modal Inconsistency

Wenfang Yao, Kejing Yin, William K. Cheung, Jia Liu, Jing Qin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by AAAI-24

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16545 2024-10-23 cs.CV 79%

PlaneSAM: Multimodal Plane Instance Segmentation Using the Segment Anything Model

Zhongchen Deng, Zhechen Yang, Chi Chen, Cheng Zeng, Yan Meng, Bisheng Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments submitted to Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14584 2024-10-21 cs.AI 79%

MCSFF: Multi-modal Consistency and Specificity Fusion Framework for Entity Alignment

Wei Ai, Wen Deng, Hongyi Chen, Jiayi Du, Tao Meng, Yuntao Shou

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments 6 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18327 2024-10-17 cs.CV 79%

MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition

Peihao Xiang, Chaohao Lin, Kaida Wu, Ou Bai

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Camera-ready Version, Accepted by ICPRS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10407 2024-10-15 cs.CL 79%

MMCFND: Multimodal Multilingual Caption-aware Fake News Detection for Low-resource Indic Languages

Shubhi Bansal, Nishit Sushil Singh, Shahid Shafi Dar, Nagendra Kumar

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08879 2024-10-14 cs.CV 79%

Multi-modal Fusion based Q-distribution Prediction for Controlled Nuclear Fusion

Shiao Wang, Yifeng Wang, Qingchuan Ma, Xiao Wang, Ning Yan, Qingquan Yang, Guosheng Xu, Jin Tang

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08739 2024-10-14 cs.CV cs.SY eess.SY 79%

MMLF: Multi-modal Multi-class Late Fusion for Object Detection with Uncertainty Estimation

Qihang Yang, Yang Zhao, Hong Cheng

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08243 2024-10-14 cs.LG cs.AI 79%

Self-Attention Mechanism in Multimodal Context for Banking Transaction Flow

Cyrile Delestre, Yoann Sola

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06395 2024-10-10 cs.LG cs.AI 79%

Multimodal Representation Learning using Adaptive Graph Construction

Weichen Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04939 2024-10-08 cs.CV 79%

PRFusion: Toward Effective and Robust Multi-Modal Place Recognition with Image and Point Cloud Fusion

Sijie Wang, Qiyu Kang, Rui She, Kai Zhao, Yang Song, Wee Peng Tay

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments accepted by IEEE TITS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01529 2024-10-03 cs.RO cs.CV 79%

Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning

Jianxiong Li, Zhihao Wang, Jinliang Zheng, Xiaoai Zhou, Guanming Wang, Guanglu Song, Yu Liu, Jingjing Liu, Ya-Qin Zhang, Junzhi Yu, Xianyuan Zhan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00152 2024-10-02 eess.IV cs.CV cs.LG q-bio.QM 79%

Multimodal Alignment of Histopathological Images Using Cell Segmentation and Point Set Matching for Integrative Cancer Analysis

Jun Jiang, Raymond Moore, Brenna Novotny, Leo Liu, Zachary Fogarty, Ray Guo, Markovic Svetomir, Chen Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments initial version

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20179 2024-10-01 eess.IV cs.CV 79%

Survival Prediction in Lung Cancer through Multi-Modal Representation Learning

Aiman Farooq, Deepak Mishra, Santanu Chaudhury

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted in WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08640 2024-09-26 cs.CL 79%

Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning

Wenwen Zhuang, Xin Huang, Xiantao Zhang, Jin Zeng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏