arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2407.09157 2024-07-15 cs.IR cs.AI cs.LG 83%

Movie Recommendation with Poster Attention via Multi-modal Transformer Feature Fusion

Linhan Xia, Yicheng Yang, Ziou Chen, Zheng Yang, Shengxin Zhu

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19654 2024-05-31 cs.AI 83%

Unlocking the Power of Spatial and Temporal Information in Medical Multimodal Pre-training

Jinxia Yang, Bing Su, Wayne Xin Zhao, Ji-Rong Wen

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.AI

Comments Accepted at ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03638 2024-05-21 cs.MM cs.HC 83%

Physical-aware Cross-modal Adversarial Network for Wearable Sensor-based Human Action Recognition

Jianyuan Ni, Hao Tang, Anne H. H. Ngu, Gaowen Liu, Yan Yan

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.MM

Comments We will be making some significant changes to the paper, including the title and methodology. We therefore wish to withdraw the paper for now

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09010 2024-04-16 cs.CV cs.LG 83%

MMA-DFER: MultiModal Adaptation of unimodal models for Dynamic Facial Expression Recognition in-the-wild

Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments accepted to CVPR 2024 ABAW Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08949 2024-04-16 cs.CL 83%

Multimodal Cross-Document Event Coreference Resolution Using Linear Semantic Transfer and Mixed-Modality Ensembles

Abhijnan Nath, Huma Jamil, Shafiuddin Rehan Ahmed, George Baker, Rahul Ghosh, James H. Martin, Nathaniel Blanchard, Nikhil Krishnaswamy

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments To appear at LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07868 2024-04-12 cs.CV 83%

MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval

Xiaojie Jin, Bowen Zhang, Weibo Gong, Kai Xu, XueQing Deng, Peng Wang, Zhao Zhang, Xiaohui Shen, Jiashi Feng

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03413 2024-04-05 cs.CV 83%

MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Deyao Zhu, Jian Ding, Mohamed Elhoseiny

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments 6 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19258 2024-03-01 cs.CV 83%

MaskFi: Unsupervised Learning of WiFi and Vision Representations for Multimodal Human Activity Recognition

Jianfei Yang, Shijie Tang, Yuecong Xu, Yunjiao Zhou, Lihua Xie

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00773 2024-02-09 cs.CV 83%

FakeOut: Leveraging Out-of-domain Self-supervision for Multi-modal Video Deepfake Detection

Gil Knafo, Ohad Fried

专题命中 视频多模态 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17797 2024-02-01 cs.CV 83%

M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval

Xingning Dong, Zipeng Feng, Chunluan Zhou, Xuzheng Yu, Ming Yang, Qingpei Guo

专题命中 视频多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01674 2024-01-04 cs.CV 83%

Transformer RGBT Tracking with Spatio-Temporal Multimodal Tokens

Dengdi Sun, Yajie Pan, Andong Lu, Chenglong Li, Bin Luo

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12721 2023-12-21 cs.CV 83%

Cross-Modal Reasoning with Event Correlation for Video Question Answering

Chengxiang Yin, Zhengping Che, Kun Wu, Zhiyuan Xu, Qinru Qiu, Jian Tang

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08083 2023-12-14 cs.CV 83%

VCD: Visual Causality Discovery for Cross-Modal Question Reasoning

Yang Liu, Ying Tan, Jingzhou Luo, Weixing Chen

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 12 pages, 6 figures. arXiv admin note: substantial text overlap with arXiv:2207.12647

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15915 2023-09-29 cs.CV 83%

Zero-Shot and Few-Shot Video Question Answering with Multi-Modal Prompts

Deniz Engin, Yannis Avrithis

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments ICCV2023 CLVL Workshop (Oral). Project page: https://engindeniz.github.io/vitis

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15313 2023-09-28 cs.CV 83%

M$^{3}$3D: Learning 3D priors using Multi-Modal Masked Autoencoders for 2D image and video understanding

Muhammad Abdullah Jamal, Omid Mohareri

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15082 2023-09-27 cs.CV 83%

RPEFlow: Multimodal Fusion of RGB-PointCloud-Event for Joint Optical Flow and Scene Flow Estimation

Zhexiong Wan, Yuxin Mao, Jing Zhang, Yuchao Dai

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICCV 2023. Project page: https://npucvr.github.io/RPEFlow Code: https://github.com/danqu130/RPEFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.02041 2023-09-06 cs.CV 83%

Learning Cross-Modal Affinity for Referring Video Object Segmentation Targeting Limited Samples

Guanghui Li, Mingqi Gao, Heng Liu, Xiantong Zhen, Feng Zheng

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12244 2023-08-24 cs.CV 83%

Multimodal Channel-Mixing: Channel and Spatial Masked AutoEncoder on Facial Action Unit Detection

Xiang Zhang, Huiyuan Yang, Taoyue Wang, Xiaotian Li, Lijun Yin

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09775 2023-08-22 cs.CV 83%

Long-range Multimodal Pretraining for Movie Understanding

Dawit Mureja Argaw, Joon-Young Lee, Markus Woodson, In So Kweon, Fabian Caba Heilbron

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14652 2023-06-01 cs.CL 83%

Denoising Bottleneck with Mutual Information Maximization for Video Multimodal Fusion

Shaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu, Binghuai Lin, Yunbo Cao, Zhifang Sui

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

Comments Accept at ACL2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10842 2023-05-09 cs.CV cs.LG cs.RO 83%

MMRNet: Improving Reliability for Multimodal Object Detection and Segmentation for Bin Picking via Multimodal Redundancy

Yuhao Chen, Hayden Gunraj, E. Zhixuan Zeng, Robbie Meyer, Maximilian Gilles, Alexander Wong

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted to CVPR TCV Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05419 2023-05-01 cs.CV 83%

Multimodal Graph Learning for Deepfake Detection

Zhiyuan Yan, Peng Sun, Yubo Lang, Shuo Du, Shanzhuo Zhang, Wei Wang, Lei Liu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17285 2023-03-31 cs.CV 83%

Decomposed Cross-modal Distillation for RGB-based Temporal Action Detection

Pilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee, Hyeran Byun

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14905 2023-03-28 cs.CV cs.AI cs.CL cs.LG cs.MM 83%

Multi-Modal Few-Shot Temporal Action Detection

Sauradip Nag, Mengmeng Xu, Xiatian Zhu, Juan-Manuel Perez-Rua, Bernard Ghanem, Yi-Zhe Song, Tao Xiang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10826 2023-03-28 cs.CV 83%

Visual Prompt Multi-Modal Tracking

Jiawen Zhu, Simiao Lai, Xin Chen, Dong Wang, Huchuan Lu

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00040 2023-03-27 cs.CV 83%

Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-Training

Dezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin, Yang Liu

专题命中 视频多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

Comments CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.10973 2022-12-05 cs.MM 83%

FakeSV: A Multimodal Benchmark with Rich Social Context for Fake News Detection on Short Video Platforms

Peng Qi, Yuyan Bu, Juan Cao, Wei Ji, Ruihao Shui, Junbin Xiao, Danding Wang, Tat-Seng Chua

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments To appear in AAAI 2023 AISI track. This version contains appendix with additional details

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12833 2022-07-27 cs.CV 83%

Multimodal-GuideNet: Gaze-Probe Bidirectional Guidance in Obstetric Ultrasound Scanning

Qianhui Men, Clare Teng, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Early accepted by MICCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03014 2022-03-29 cs.CV 83%

Learnable Irrelevant Modality Dropout for Multimodal Action Recognition on Modality-Specific Annotated Videos

Saghir Alfasly, Jian Lu, Chen Xu, Yuru Zou

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.09322 2021-11-16 cs.CV 83%

MM-ViT: Multi-Modal Video Transformer for Compressed Video Action Recognition

Jiawei Chen, Chiu Man Ho

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Winter Conference on Applications of Computer Vision (WACV) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏