arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4728 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4728 篇

2308.12049 2023-08-24 cs.CV cs.AI 62%

Towards Privacy-Supporting Fall Detection via Deep Unsupervised RGB2Depth Adaptation

Hejun Xiao, Kunyu Peng, Xiangsheng Huang, Alina Roitberg1, Hao Li, Zhaohui Wang, Rainer Stiefelhagen

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11093 2023-08-23 cs.CV cs.AI cs.LG 62%

Video OWL-ViT: Temporally-consistent open-world localization in video

Georg Heigold, Matthias Minderer, Alexey Gritsenko, Alex Bewley, Daniel Keysers, Mario Lučić, Fisher Yu, Thomas Kipf

专题命中 视频多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06897 2023-08-15 cs.CV cs.MM 62%

Orthogonal Temporal Interpolation for Zero-Shot Video Recognition

Yan Zhu, Junbao Zhuo, Bin Ma, Jiajia Geng, Xiaoming Wei, Xiaolin Wei, Shuhui Wang

专题命中 视频多模态 :image-text(abstract);分类 cs.CV、cs.MM

Journal ref Proceedings of the 31st ACM International Conference on Multimedia (MM '23), October 29-November 3, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06571 2023-08-15 cs.CV cs.AI 62%

ModelScope Text-to-Video Technical Report

Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, Shiwei Zhang

专题命中 视频多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments Technical report. Project page: \url{https://modelscope.cn/models/damo/text-to-video-synthesis/summary}

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08966 2023-08-14 cs.MM cs.CV 62%

Training Multimedia Event Extraction With Generated Images and Captions

Zilin Du, Yunxin Li, Xu Guo, Yidan Sun, Boyang Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04047 2023-08-09 cs.CV cs.AI cs.RO 62%

SODFormer: Streaming Object Detection with Transformer Using Events and Frames

Dianze Li, Jianing Li, Yonghong Tian

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 18 pages, 15 figures, in IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03908 2023-08-09 cs.CV cs.AI cs.LG 62%

ViLP: Knowledge Exploration using Vision, Language, and Pose Embeddings for Video Action Recognition

Soumyabrata Chaudhuri, Saumik Bhattacharya

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 7 pages, 3 figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03267 2023-08-08 cs.CV cs.AI 62%

Redundancy-aware Transformer for Video Question Answering

Yicong Li, Xun Yang, An Zhang, Chun Feng, Xiang Wang, Tat-Seng Chua

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ACM MM23

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.11589 2023-08-04 cs.CV cs.AI cs.LG 62%

SBNet: Segmentation-based Network for Natural Language-based Vehicle Search

Sangrok Lee, Taekang Woo, Sang Hun Lee

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 7 pages, 4 figures, CVPR Workshop Paper

Journal ref 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 4049-4055

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02159 2023-07-19 cs.CV cs.MM 62%

Robustness Analysis of Video-Language Models Against Visual and Language Perturbations

Madeline C. Schiappa, Shruti Vyas, Hamid Palangi, Yogesh S. Rawat, Vibhav Vineet

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments NeurIPS 2022 Datasets and Benchmarks Track. This projects webpage is located at https://bit.ly/3CNOly4

Journal ref Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03153 2023-07-07 cs.IR cs.CV cs.MM 62%

MultiVENT: Multilingual Videos of Events with Aligned Natural Text

Kate Sanders, David Etter, Reno Kriz, Benjamin Van Durme

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01947 2023-07-06 cs.CV cs.AI cs.IR 62%

Causal Video Summarizer for Video Exploration

Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen, Andrew Brown, Marcel Worring

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments This paper is accepted by IEEE International Conference on Multimedia and Expo (ICME), 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01378 2023-07-06 cs.CV cs.AI eess.IV 62%

A CNN regression model to estimate buildings height maps using Sentinel-1 SAR and Sentinel-2 MSI time series

Ritu Yadav, Andrea Nascetti, Yifang Ban

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15255 2023-06-28 cs.CV cs.CL 62%

GroundNLQ @ Ego4D Natural Language Queries Challenge 2023

Zhijian Hou, Lei Ji, Difei Gao, Wanjun Zhong, Kun Yan, Chao Li, Wing-Kwong Chan, Chong-Wah Ngo, Nan Duan, Mike Zheng Shou

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments 5 pages, 2 figures, 4 tables, the champion solution for Ego4D Natural Language Queries Challenge in CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14412 2023-06-27 cs.CV cs.MM 62%

A Solution to CVPR'2023 AQTC Challenge: Video Alignment for Multi-Step Inference

Chao Zhang, Shiwei Wu, Sirui Zhao, Tong Xu, Enhong Chen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Comments 5 pages, 1 figure, technical report for track3 of CVPR 2023 LOVEU challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11025 2023-06-21 cs.LG cs.AI cs.CL q-fin.ST 62%

Temporal Data Meets LLM -- Explainable Financial Time Series Forecasting

Xinli Yu, Zheng Chen, Yuan Ling, Shujing Dong, Zongyi Liu, Yanbin Lu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07187 2023-06-13 cs.MM cs.IR cs.LG cs.SD eess.AS 62%

Video-to-Music Recommendation using Temporal Alignment of Segments

Laure Prétet, Gaël Richard, Clément Souchier, Geoffroy Peeters

专题命中 视频多模态 :cross-modal(abstract);分类 cs.MM、eess.AS

Journal ref IEEE Transactions on Multimedia, 18 February 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03517 2023-06-06 cs.CL cs.AI 62%

Few-shot Domain-Adaptive Visually-fused Event Detection from Text

Farhad Moghimifar, Fatemeh Shiri, Van Nguyen, Reza Haffari, Yuan-Fang Li

专题命中 视频多模态 :image-text(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10918 2023-06-01 cs.CV cs.CL cs.IR 62%

CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding

Zhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao, Kun Yan, Wing-Kwong Chan, Chong-Wah Ngo, Zheng Shou, Nan Duan

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments ACL 2023 Camera Ready. 14 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12218 2023-05-23 cs.CV cs.AI cs.IR 62%

Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment

Peng Jin, Hao Li, Zesen Cheng, Jinfa Huang, Zhennan Wang, Li Yuan, Chang Liu, Jie Chen

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments IJCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03204 2023-05-08 cs.CV cs.CL 62%

VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation

Xilun Chen, Lili Yu, Wenhan Xiong, Barlas Oğuz, Yashar Mehdad, Wen-tau Yih

专题命中 视频多模态 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13143 2023-04-27 cs.AI cs.CV cs.LG 62%

Self-Supervised Temporal Analysis of Spatiotemporal Data

Yi Cao, Swetava Ganguli, Vipul Pandey

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted for oral presentation at the 43rd IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2023, Pasadena, California. 4 pages and 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.12561 2023-04-26 cs.CV cs.MM 62%

TCR: Short Video Title Generation and Cover Selection with Attention Refinement

Yakun Yu, Jiuding Yang, Weidong Guo, Hui Liu, Yu Xu, Di Niu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Comments Accepted by PAKDD23

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09751 2023-04-20 cs.CV cs.AI cs.LG 62%

Skeleton-based action analysis for ADHD diagnosis

Yichun Li, Yi Li, Rajesh Nair, Syed Mohsen Naqvi

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05038 2023-04-20 cs.CL cs.CV 62%

Fighting FIRe with FIRE: Assessing the Validity of Text-to-Video Retrieval Benchmarks

Pedro Rodriguez, Mahmoud Azab, Becka Silvert, Renato Sanchez, Linzy Labson, Hardik Shah, Seungwhan Moon

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments EACL 2023 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05711 2023-04-06 cs.CV cs.CL cs.LG 62%

Synopses of Movie Narratives: a Video-Language Dataset for Story Understanding

Yidan Sun, Qin Chao, Yangfeng Ji, Boyang Li

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 25 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13916 2023-03-31 cs.CV cs.AI cs.LG 62%

Towards Good Practices for Missing Modality Robust Action Recognition

Sangmin Woo, Sumin Lee, Yeonju Park, Muhammad Adi Nugroho, Changick Kim

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments AAAI 2023 (Oral); Code: https://github.com/sangminwoo/ActionMAE

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05093 2023-03-10 cs.CV cs.CL 62%

Improving Video Retrieval by Adaptive Margin

Feng He, Qi Wang, Zhifan Feng, Wenbin Jiang, Yajuan Lv, Yong zhu, Xiao Tan

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by SIGIR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03353 2023-02-23 cs.CL cs.AI cs.NE cs.RO 62%

Learning Bidirectional Action-Language Translation with Limited Supervision and Incongruent Input

Ozan Özdemir, Matthias Kerzel, Cornelius Weber, Jae Hee Lee, Muhammad Burhan Hafez, Patrick Bruns, Stefan Wermter

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Published in: Applied Artificial Intelligence, 37:1, 2179167

Journal ref Applied Artificial Intelligence Volume 37, 2023 - Issue 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11507 2023-01-30 cs.CV cs.CL cs.LG 62%

Semi-Parametric Video-Grounded Text Generation

Sungdong Kim, Jin-Hwa Kim, Jiyoung Lee, Minjoon Seo

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Preprint (16 pages, 5 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏