arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2309.09067 2023-09-20 cs.CV 79%

MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision Transformer

Fudong Lin, Summer Crawford, Kaleb Guillot, Yihe Zhang, Yan Chen, Xu Yuan, Li Chen, Shelby Williams, Robert Minvielle, Xiangming Xiao, Drew Gholson, Nicolas Ashwell, Tri Setiyono, Brenda Tubana, Lu Peng, Magdy Bayoumi, Nian-Feng Tzeng

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Journal ref ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.02320 2023-09-06 physics.geo-ph cs.AI cs.LG 79%

SeisCLIP: A seismology foundation model pre-trained by multi-modal data for multi-purpose seismic feature extraction

Xu Si, Xinming Wu, Hanlin Sheng, Jun Zhu, Zefeng Li

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments 27 pages, 9 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04129 2023-09-06 cs.CV 79%

Cross-modal Orthogonal High-rank Augmentation for RGB-Event Transformer-trackers

Zhiyu Zhu, Junhui Hou, Dapeng Oliver Wu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments accepted by ICCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14383 2023-08-29 cs.CV 79%

Multi-Modal Neural Radiance Field for Monocular Dense SLAM with a Light-Weight ToF Sensor

Xinyang Liu, Yijin Li, Yanbin Teng, Hujun Bao, Guofeng Zhang, Yinda Zhang, Zhaopeng Cui

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2023 (Oral). Project Page: https://zju3dv.github.io/tof_slam/

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07205 2023-08-29 cs.CV 79%

Multimodal Motion Conditioned Diffusion Model for Skeleton-based Video Anomaly Detection

Alessandro Flaborea, Luca Collorone, Guido D'Amely, Stefano D'Arrigo, Bardh Prenkaj, Fabio Galasso

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11185 2023-08-23 cs.CV 79%

MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation

Najmeh Sadoughi, Xinyu Li, Avijit Vajpayee, David Fan, Bing Shuai, Hector Santos-Villalobos, Vimal Bhat, Rohith MV

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2023 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05638 2023-08-21 cs.CV 79%

Cross-Modal Learning with 3D Deformable Attention for Action Recognition

Sangwon Kim, Dasom Ahn, Byoung Chul Ko

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02459 2023-08-07 cs.RO cs.AI cs.LG 79%

Nonprehensile Planar Manipulation through Reinforcement Learning with Multimodal Categorical Exploration

Juan Del Aguila Ferrandis, João Moura, Sethu Vijayakumar

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03906 2023-07-31 cs.CV 79%

Cross-modal Manifold Cutmix for Self-supervised Video Representation Learning

Srijan Das, Michael S. Ryoo

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at MVA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10346 2023-07-21 cs.HC cs.MM 79%

Estudio de la Experiencia de Usuario mediante un Sistema de Dashboards de Análisis de Aprendizaje Multimodal

Álvaro Becerra, Roberto Daza, Ruth Cobos, Aythami Morales, Julian Fierrez

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted in "XXIII CONGRESO INTERNACIONAL DE INTERACCIÓN PERSONA-ORDENADOR 2023". Article in Spanish language. The abstract in English and Spanish. There is an extended abstract of 2 pages in English

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07807 2023-07-18 eess.IV cs.CV 79%

MUVF-YOLOX: A Multi-modal Ultrasound Video Fusion Network for Renal Tumor Diagnosis

Junyu Li, Han Huang, Dong Ni, Wufeng Xue, Dongmei Zhu, Jun Cheng

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01047 2023-07-04 cs.CV 79%

Cross-modal Place Recognition in Image Databases using Event-based Sensors

Xiang Ji, Jiaxin Wei, Yifu Wang, Huiliang Shang, Laurent Kneip

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00226 2023-07-04 cs.CV cs.LG 79%

S-Omninet: Structured Data Enhanced Universal Multimodal Learning Architecture

Ye Xue, Diego Klabjan, Jean Utke

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14392 2023-06-27 cs.CV 79%

ContentCTR: Frame-level Live Streaming Click-Through Rate Prediction with Multimodal Transformer

Jiaxin Deng, Dong Shen, Shiyao Wang, Xiangyu Wu, Fan Yang, Guorui Zhou, Gaofeng Meng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12156 2023-06-07 cs.LG cs.AI 79%

Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling

Xinlu Zhang, Shiyang Li, Zhiyu Chen, Xifeng Yan, Linda Petzold

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02845 2023-06-06 cs.AI 79%

Interpretable Multimodal Emotion Recognition using Facial Features and Physiological Signals

Puneet Kumar, Xiaobai Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted for Oral Presentation in DAI 2023 (https://rbcdsai.iitm.ac.in/DAI-2023/program.html)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00640 2023-06-02 cs.CV eess.IV 79%

Multi-Modal Deep Learning for Multi-Temporal Urban Mapping With a Partly Missing Optical Modality

Sebastian Hafner, Yifang Ban

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 4 pages, 2 figures, accepted for publication in the IGARSS 2023 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19624 2023-06-01 cs.CV 79%

A Multi-Modal Transformer Network for Action Detection

Matthew Korban, Scott T. Acton, Peter Youngs

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Journal ref Pattern Recognition 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08352 2023-05-31 cs.CV 79%

MHSCNet: A Multimodal Hierarchical Shot-aware Convolutional Network for Video Summarization

Wujiang Xu, Runzhong Wang, Xiaobo Guo, Shaoshuai Li, Qiongxu Ma, Yunan Zhao, Sheng Guo, Zhenfeng Zhu, Junchi Yan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12448 2023-05-26 cs.CV 79%

CMD: Self-supervised 3D Action Representation Learning with Cross-modal Mutual Distillation

Yunyao Mao, Wengang Zhou, Zhenbo Lu, Jiajun Deng, Houqiang Li

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments To appear in ECCV 2022 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12561 2023-05-23 cs.HC cs.CV 79%

M2LADS: A System for Generating MultiModal Learning Analytics Dashboards in Open Education

Álvaro Becerra, Roberto Daza, Ruth Cobos, Aythami Morales, Mutlu Cukurova, Julian Fierrez

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted in "Workshop on Open Education Resources (OER) of COMPSAC 2023"

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08401 2023-05-18 cs.CV cs.LG 79%

Multimodal Short Video Rumor Detection System Based on Contrastive Learning

Yuxing Yang, Junhao Zhao, Siyi Wang, Xiangyu Min, Pengchao Wang, Haizhou Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14407 2023-05-02 cs.CV 79%

ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Junke Wang, Dongdong Chen, Chong Luo, Xiyang Dai, Lu Yuan, Zuxuan Wu, Yu-Gang Jiang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10590 2023-04-19 cs.CV 79%

Multi-modal Facial Action Unit Detection with Large Pre-trained Models for the 5th Competition on Affective Behavior Analysis in-the-wild

Yufeng Yin, Minh Tran, Di Chang, Xinrui Wang, Mohammad Soleymani

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 8 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10849 2023-04-12 cs.CV 79%

Multi-modal Facial Affective Analysis based on Masked Autoencoder

Wei Zhang, Bowen Ma, Feng Qiu, Yu Ding

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.11732 2023-03-22 cs.CV 79%

Multi-modal Prompting for Low-Shot Temporal Action Localization

Chen Ju, Zeqian Li, Peisen Zhao, Ya Zhang, Xiaopeng Zhang, Qi Tian, Yanfeng Wang, Weidi Xie

专题命中 视频多模态 :multi-modal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02136 2023-03-07 cs.CV 79%

Efficient End-to-End Video Question Answering with Pyramidal Multimodal Transformer

Min Peng, Chongyang Wang, Yu Shi, Xiang-Dong Zhou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06031 2023-03-03 cs.CV 79%

Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive Learning

Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu, Huan Yang, Jianlong Fu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10465 2023-02-22 cs.CV 79%

A Flexible Multi-view Multi-modal Imaging System for Outdoor Scenes

Meng Zhang, Wenxuan Guo, Bohao Fan, Yifan Chen, Jianjiang Feng, Jie Zhou

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11314 2023-02-22 cs.CV 79%

Modality Mixer for Multi-modal Action Recognition

Sumin Lee, Sangmin Woo, Yeonju Park, Muhammad Adi Nugroho, Changick Kim

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏