arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2411.12829 2024-11-21 cs.HC cs.CL cs.RO 74%

Human-Robot Dialogue Annotation for Multi-Modal Common Ground

Claire Bonial, Stephanie M. Lukin, Mitchell Abrams, Anthony Baker, Lucia Donatelli, Ashley Foots, Cory J. Hayes, Cassidy Henry, Taylor Hudson, Matthew Marge, Kimberly A. Pollard, Ron Artstein, David Traum, Clare R. Voss

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CL

Comments 52 pages, 14 figures

Journal ref Language Resources and Evaluation 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04963 2024-11-08 cs.CV 74%

VAIR: Visuo-Acoustic Implicit Representations for Low-Cost, Multi-Modal Transparent Surface Reconstruction in Indoor Scenes

Advaith V. Sethuraman, Onur Bagoren, Harikrishnan Seetharaman, Dalton Richardson, Joseph Taylor, Katherine A. Skinner

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments https://umfieldrobotics.github.io/VAIR_site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12592 2024-10-17 cs.CV cs.LG 74%

Cocoon: Robust Multi-Modal Perception with Uncertainty-Aware Sensor Fusion

Minkyoung Cho, Yulong Cao, Jiachen Sun, Qingzhao Zhang, Marco Pavone, Jeong Joon Park, Heng Yang, Z. Morley Mao

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07673 2024-10-11 cs.LG cs.AI 74%

Multimodal Clickbait Detection by De-confounding Biases Using Causal Representation Inference

Jianxing Yu, Shiqi Wang, Han Yin, Zhenlong Sun, Ruobing Xie, Bo Zhang, Yanghui Rao

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00469 2024-10-02 cs.CV 74%

Deep Multimodal Fusion for Semantic Segmentation of Remote Sensing Earth Observation Data

Ivica Dimitrovski, Vlatko Spasev, Ivan Kitanovski

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19305 2024-10-01 cs.CV eess.IV 74%

EEPNet: Efficient Edge Pixel-based Matching Network for Cross-Modal Dynamic Registration between LiDAR and Camera

Yuanchao Yue, Hui Yuan, Suai Li, Qi Jiang

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15801 2024-09-25 cs.CV 74%

DIAL: Dense Image-text ALignment for Weakly Supervised Semantic Segmentation

Soojin Jang, Jungmin Yun, Junehyoung Kwon, Eunju Lee, Youngbin Kim

专题命中 多模态训练与对齐 :image-text(title);分类 cs.CV

Comments accepted by the European Conference on Computer Vision (ECCV), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10499 2024-08-21 cs.HC cs.AI cs.PL 74%

ProgramAlly: Creating Custom Visual Access Programs via Multi-Modal End-User Programming

Jaylin Herskovitz, Andi Xu, Rahaf Alharbi, Anhong Guo

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.AI

Comments UIST 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08431 2024-08-19 cs.AI 74%

Multi-Modal Dialogue State Tracking for Playing GuessWhich Game

Wei Pang, Ruixue Duan, Jinfu Yang, Ning Li

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.AI

Comments Published at CICAI 2023 (CAAI-A), codes at https://github.com/xubuvd/GuessWhich

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18362 2024-07-29 eess.IV cs.CV cs.LG 74%

Retinal IPA: Iterative KeyPoints Alignment for Multimodal Retinal Imaging

Jiacheng Wang, Hao Li, Dewei Hu, Rui Xu, Xing Yao, Yuankai K. Tao, Ipek Oguz

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17726 2024-07-26 cs.LG cs.CV 74%

Multi-modal Data Binding for Survival Analysis Modeling with Incomplete Data and Annotations

Linhao Qu, Dan Huang, Shaoting Zhang, Xiaosong Wang

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments Accepted by MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14812 2024-07-23 cs.CV 74%

GaitMA: Pose-guided Multi-modal Feature Fusion for Gait Recognition

Fanxu Min, Shaoxiang Guo, Fan Hao, Junyu Dong

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments Accepted to ICME 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17469 2024-06-26 cs.CV 74%

Cross-Modal Spherical Aggregation for Weakly Supervised Remote Sensing Shadow Removal

Kaichen Chi, Wei Jing, Junjie Li, Qiang Li, Qi Wang

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

Comments 9pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04249 2024-06-07 cs.CV 74%

Conv-INR: Convolutional Implicit Neural Representation for Multimodal Visual Signals

Zhicheng Cai

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10254 2024-05-24 eess.IV cs.CV cs.LG 74%

PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology

George Shaikovski, Adam Casson, Kristen Severson, Eric Zimmermann, Yi Kan Wang, Jeremy D. Kunz, Juan A. Retamero, Gerard Oakley, David Klimstra, Christopher Kanan, Matthew Hanna, Michal Zelechowski, Julian Viret, Neil Tenenholtz, James Hall, Nicolo Fusi, Razik Yousfi, Peter Hamilton, William A. Moye, Eugene Vorontsov, Siqi Liu, Thomas J. Fuchs

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07573 2024-05-14 cs.CV 74%

MaskFuser: Masked Fusion of Joint Multi-Modal Tokenization for End-to-End Autonomous Driving

Yiqun Duan, Xianda Guo, Zheng Zhu, Zhen Wang, Yu-Kai Wang, Chin-Teng Lin

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13120 2024-03-29 cs.CV 74%

Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer

Zhen Zhao, Jingqun Tang, Chunhui Lin, Binghong Wu, Can Huang, Hao Liu, Xin Tan, Zhizhong Zhang, Yuan Xie

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Comments Accepted to CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15462 2024-03-26 eess.IV cs.CV cs.LG 74%

FUELVISION: A Multimodal Data Fusion and Multimodel Ensemble Algorithm for Wildfire Fuels Mapping

Riyaaz Uddien Shaik, Mohamad Alipour, Eric Rowell, Bharathan Balaji, Adam Watts, Ertugrul Taciroglu

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments 40 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06059 2024-03-12 cs.CV 74%

Test-time Distribution Learning Adapter for Cross-modal Visual Reasoning

Yi Zhang, Ce Zhang

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

Comments Accepted by ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03405 2024-03-07 cs.CV 74%

Causality-based Cross-Modal Representation Learning for Vision-and-Language Navigation

Liuyi Wang, Zongtao He, Ronghao Dang, Huiyi Chen, Chengju Liu, Qijun Chen

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13418 2024-01-25 cs.CV 74%

Serial fusion of multi-modal biometric systems

Gian Luca Marcialis, Paolo Mastinu, Fabio Roli

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

Journal ref IEEE International Workshop on Biometric Measurements and Systems for Security and Medical Applications (BioMS2010), September, 9, 2010, Taranto (Italy), ISBN: 978-1-4244-6302-2

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08261 2024-01-11 cs.CV 74%

GraphAlign: Enhancing Accurate Feature Alignment by Graph matching for Multi-Modal 3D Object Detection

Ziying Song, Haiyue Wei, Lin Bai, Lei Yang, Caiyan Jia

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01592 2024-01-10 cs.CL 74%

Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment

Cong-Duy Nguyen, The-Anh Vu-Le, Thong Nguyen, Tho Quan, Luu Anh Tuan

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06855 2023-12-13 cs.LG cs.CL 74%

Multimodal Pretraining of Medical Time Series and Notes

Ryan King, Tianbao Yang, Bobak Mortazavi

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13022 2023-11-23 cs.LG cs.CV 74%

Unsupervised Multimodal Surface Registration with Geometric Deep Learning

Mohamed A. Suliman, Logan Z. J. Williams, Abdulah Fawaz, Emma C. Robinson

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12040 2023-11-22 q-bio.QM cs.AI cs.LG 74%

TransCDR: a deep learning model for enhancing the generalizability of cancer drug response prediction through transfer learning and multimodal data fusion for drug representation

Xiaoqiong Xia, Chaoyu Zhu, Yuqi Shan, Fan Zhong, Lei Liu

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI

Comments 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15109 2023-09-27 cs.CV cs.RO 74%

DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge Distillation

Zeyu Wang, Dingwen Li, Chenxu Luo, Cihang Xie, Xiaodong Yang

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.00993 2023-08-22 eess.SP cs.AI cs.HC 74%

Data Fusion in Neuromarketing: Multimodal Analysis of Biosignals, Lifecycle Stages, Current Advances, Datasets, Trends, and Challenges

Mario Quiles Pérez, Enrique Tomás Martínez Beltrán, Sergio López Bernal, Eduardo Horna Prat, Luis Montesano Del Campo, Lorenzo Fernández Maimó, Alberto Huertas Celdrán

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.AI

Comments 26 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08414 2023-08-17 cs.CV 74%

Tem-adapter: Adapting Image-Text Pretraining for Video Question Answer

Guangyi Chen, Xiao Liu, Guangrun Wang, Kun Zhang, Philip H. S. Torr, Xiao-Ping Zhang, Yansong Tang

专题命中 多模态训练与对齐 :image-text(title);分类 cs.CV

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15301 2023-07-31 cs.CV 74%

Attentive Multimodal Fusion for Optical and Scene Flow

Youjie Zhou, Guofeng Mei, Yiming Wang, Fabio Poiesi, Yi Wan

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

Comments This work is accepted for publication in IEEE Robotics and Automation Letters

详情

展开后加载摘要…

URL PDF HTML 收藏