arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2302.00912 2023-02-08 cs.CV 79%

Advances and Challenges in Multimodal Remote Sensing Image Registration

Bai Zhu, Liang Zhou, Simiao Pu, Jianwei Fan, Yuanxin Ye

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.04780 2023-01-20 cs.CV 79%

MAiVAR: Multimodal Audio-Image and Video Action Recognizer

Muhammad Bilal Shaikh, Douglas Chai, Syed Mohammed Shamsul Islam, Naveed Akhtar

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Peer reviewed & accepted at IEEE VCIP 2022 (http://www.vcip2022.org/)

Journal ref 2022 IEEE International Conference on Visual Communications and Image Processing (VCIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00167 2023-01-13 cs.CV 79%

Event-Based Fusion for Motion Deblurring with Cross-modal Attention

Lei Sun, Christos Sakaridis, Jingyun Liang, Qi Jiang, Kailun Yang, Peng Sun, Yaozu Ye, Kaiwei Wang, Luc Van Gool

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ECCV 2022 as oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10596 2023-01-12 cs.CV 79%

Open-Vocabulary Temporal Action Detection with Off-the-Shelf Image-Text Features

Vivek Rathod, Bryan Seybold, Sudheendra Vijayanarasimhan, Austin Myers, Xiuye Gu, Vighnesh Birodkar, David A. Ross

专题命中 视频多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02177 2023-01-05 cs.LG cs.CL 79%

GCNet: Graph Completion Network for Incomplete Multimodal Learning in Conversation

Zheng Lian, Lan Chen, Licai Sun, Bin Liu, Jianhua Tao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.14143 2023-01-02 cs.CV 79%

Multimodal Wildland Fire Smoke Detection

Siddhant Baldota, Shreyas Anantha Ramaprasad, Jaspreet Kaur Bhamra, Shane Luna, Ravi Ramachandra, Eugene Zen, Harrison Kim, Daniel Crawl, Ismael Perez, Ilkay Altintas, Garrison W. Cottrell, Mai H. Nguyen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.06345 2022-12-27 cs.MM 79%

Self-supervised Multi-Modal Video Forgery Attack Detection

Chenhui Zhao, Xiang Li, Rabih Younes

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09522 2022-12-20 cs.CV 79%

MIST: Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question Answering

Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yi Yang, Mike Zheng Shou

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08859 2022-12-20 cs.RO cs.CV cs.LG 79%

iCub! Do you recognize what I am doing?: multimodal human action recognition on multisensory-enabled iCub robot

Kas Kniesmeijer, Murat Kirtay

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 5 figures and 1 table. International Conference on Social Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04700 2022-12-12 cs.CV 79%

Tencent AVS: A Holistic Ads Video Dataset for Multi-modal Scene Segmentation

Jie Jiang, Zhimin Li, Jiangfeng Xiong, Rongwei Quan, Qinglin Lu, Wei Liu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13804 2022-11-28 cs.HC cs.CL 79%

On the Linguistic and Computational Requirements for Creating Face-to-Face Multimodal Human-Machine Interaction

João Ranhel, Cacilda Vilela de Lima

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03785 2022-11-08 cs.AI cs.RO 79%

Learning Visual Locomotion with Cross-Modal Supervision

Antonio Loquercio, Ashish Kumar, Jitendra Malik

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.AI

Comments Learning to walk from pixels in the real world by using proprioception as supervision. Project page for videos and code: https://antonilo.github.io/vision_locomotion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.06125 2022-11-04 cs.LG cs.AI physics.ao-ph 79%

Hurricane Forecasting: A Novel Multimodal Machine Learning Framework

Léonard Boussioux, Cynthia Zeng, Théo Guénais, Dimitris Bertsimas

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Published by the AMS' Weather and Forecasting journal; Spotlight talk at NeurIPS 2021, Tackling Climate Change with AI ; https://journals.ametsoc.org/view/journals/wefo/37/6/WAF-D-21-0091.1.xml

Journal ref 2022, Weather and Forecasting, 37(6), 817-831

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14512 2022-10-27 cs.CV 79%

End-to-End Multimodal Representation Learning for Video Dialog

Huda Alamri, Anthony Bilic, Michael Hu, Apoorva Beedu, Irfan Essa

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12649 2022-10-25 cs.CV cs.RO 79%

Anticipative Feature Fusion Transformer for Multi-Modal Action Anticipation

Zeyun Zhong, David Schneider, Michael Voit, Rainer Stiefelhagen, Jürgen Beyerer

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to WACV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08452 2022-10-18 cs.CV 79%

Efficient Cross-Modal Video Retrieval with Meta-Optimized Frames

Ning Han, Xun Yang, Ee-Peng Lim, Hao Chen, Qianru Sun

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.07886 2022-10-17 cs.CV cs.RO 79%

PedFormer: Pedestrian Behavior Prediction via Cross-Modal Attention Modulation and Gated Multitask Learning

Amir Rasouli, Iuliia Kotseruba

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 8 pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05840 2022-10-13 cs.CV 79%

LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream Videos

Jielin Qiu, Franck Dernoncourt, Trung Bui, Zhaowen Wang, Ding Zhao, Hailin Jin

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04331 2022-10-11 cs.CV 79%

Students taught by multimodal teachers are superior action recognizers

Gorjan Radevski, Dusan Grujicic, Matthew Blaschko, Marie-Francine Moens, Tinne Tuytelaars

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Extended abstract accepted at the 2nd Ego4D Workshop @ ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13371 2022-10-07 cs.CV 79%

FitCLIP: Refining Large-Scale Pretrained Image-Text Models for Zero-Shot Video Understanding Tasks

Santiago Castro, Fabian Caba Heilbron

专题命中 视频多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted at BMVC 2022. It includes the supplementary material. The margins and page size were modified to fit the arXiv ID stamp on the left side

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.06208 2022-10-03 cs.LG cs.AI cs.HC eess.SP 79%

Identification of Cognitive Workload during Surgical Tasks with Multimodal Deep Learning

Kaizhe Jin, Adrian Rubio-Solis, Ravi Naik, Tochukwu Onyeogulu, Amirul Islam, Salman Khan, Izzeddin Teeti, James Kinross, Daniel R Leff, Fabio Cuzzolin, George Mylonas

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10170 2022-09-22 cs.CV 79%

FV2ES: A Fully End2End Multimodal System for Fast Yet Effective Video Emotion Recognition Inference

Qinglan Wei, Xuling Huang, Yuan Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02080 2022-08-04 cs.CV 79%

A Feature-space Multimodal Data Augmentation Technique for Text-video Retrieval

Alex Falcon, Giuseppe Serra, Oswald Lanz

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted for presentation at 30th ACM International Conference on Multimedia (ACM MM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01954 2022-08-04 cs.CV 79%

Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in Videos

Juncheng Li, Junlin Xie, Linchao Zhu, Long Qian, Siliang Tang, Wenqiao Zhang, Haochen Shi, Shengyu Zhang, Longhui Wei, Qi Tian, Yueting Zhuang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ACM Multimedia 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10123 2022-07-22 cs.CV 79%

Animation from Blur: Multi-modal Blur Decomposition with Motion Guidance

Zhihang Zhong, Xiao Sun, Zhirong Wu, Yinqiang Zheng, Stephen Lin, Imari Sato

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments ECCV2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.05759 2022-07-19 cs.CV 79%

Multimodal Transformer with Variable-length Memory for Vision-and-Language Navigation

Chuang Lin, Yi Jiang, Jianfei Cai, Lizhen Qu, Gholamreza Haffari, Zehuan Yuan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.06187 2022-07-12 cs.CV 79%

Calibrating Class Weights with Multi-Modal Information for Partial Video Domain Adaptation

Xiyu Wang, Yuecong Xu, Kezhi Mao, Jianfei Yang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ACM Multimedia (ACMMM) 2022, update to camera-ready version. 8 pages of text, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.03317 2022-07-08 cs.AI 79%

Multimodal Feature Extraction for Memes Sentiment Classification

Sofiane Ouaari, Tsegaye Misikir Tashu, Tomas Horvath

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02756 2022-07-07 cs.CV 79%

STVGFormer: Spatio-Temporal Video Grounding with Static-Dynamic Cross-Modal Understanding

Zihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin, Tiancai Ye, Wei-Shi Zheng

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Technical report. The 1st place solution in the HC-STVG track of the 4th Person in Context Challenge(2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01241 2022-07-05 cs.CV 79%

OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and Classification

Ye Liu, Lingfeng Qiao, Di Yin, Zhuoxuan Jiang, Xinghua Jiang, Deqiang Jiang, Bo Ren

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏