arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2301.02911 2023-01-18 cs.CV 57%

Towards early prediction of neurodevelopmental disorders: Computational model for Face Touch and Self-adaptors in Infants

Bruno Tafur, Marwa Mahmoud, Staci Weiss

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.04548 2023-01-18 cs.HC cs.CV 57%

Inconsistencies in the Definition and Annotation of Student Engagement in Virtual Learning Datasets: A Critical Review

Shehroz S. Khan, Ali Abedi, Tracey Colella

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00514 2023-01-03 cs.CV 57%

Rethinking the Video Sampling and Reasoning Strategies for Temporal Sentence Grounding

Jiahao Zhu, Daizong Liu, Pan Zhou, Xing Di, Yu Cheng, Song Yang, Wenzheng Xu, Zichuan Xu, Yao Wan, Lichao Sun, Zeyu Xiong

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by EMNLP Findings, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.14289 2023-01-03 cs.CV 57%

High-temporal-resolution event-based vehicle detection and tracking

Zaid El-Shair, Samir Rawashdeh

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 38 pages, 9 figures, 4 tables

Journal ref Optical Engineering 62(3), 031209 (22 December 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.13340 2022-12-29 cs.CV 57%

Recovering Surveillance Video Using RF Cues

Xiang Li, Rabih Younes

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.03918 2022-12-14 cs.CV 57%

Depth Quality-Inspired Feature Manipulation for Efficient RGB-D and Video Salient Object Detection

Wenbo Zhang, Keren Fu, Zhuo Wang, Ge-Peng Ji, Qijun Zhao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2107.01779

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04533 2022-12-13 cs.CV cs.LG 57%

Deep Architectures for Content Moderation and Movie Content Rating

Fatih Cagatay Akyon, Alptekin Temizel

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.03690 2022-12-08 cs.CV cs.RO 57%

Gaussian Radar Transformer for Semantic Segmentation in Noisy Radar Data

Matthias Zeller, Jens Behley, Michael Heidingsfeld, Cyrill Stachniss

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted for publication in IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.03490 2022-12-08 cs.CV 57%

SimVTP: Simple Video Text Pre-training with Masked Autoencoders

Yue Ma, Tianyu Yang, Yin Shan, Xiu Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Github: https://github.com/mayuelala/SimVTP

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06343 2022-12-07 cs.CV 57%

Change Detection Meets Visual Question Answering

Zhenghang Yuan, Lichao Mou, Zhitong Xiong, Xiaoxiang Zhu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.09760 2022-12-07 cs.CV 57%

HCMS: Hierarchical and Conditional Modality Selection for Efficient Video Recognition

Zejia Weng, Zuxuan Wu, Hengduo Li, Jingjing Chen, Yu-Gang Jiang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09210 2022-12-06 cs.HC cs.CV 57%

edBB-Demo: Biometrics and Behavior Analysis for Online Educational Platforms

Roberto Daza, Aythami Morales, Ruben Tolosana, Luis F. Gomez, Julian Fierrez, Javier Ortega-Garcia

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in "AAAI-23 Conference on Artificial Intelligence (Demonstration Program)"

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16951 2022-12-01 cs.CV 57%

Weakly Supervised 3D Multi-person Pose Estimation for Large-scale Scenes based on Monocular Camera and Single LiDAR

Peishan Cong, Yiteng Xu, Yiming Ren, Juze Zhang, Lan Xu, Jingya Wang, Jingyi Yu, Yuexin Ma

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11351 2022-11-22 cs.CV 57%

Are All Combinations Equal? Combining Textual and Visual Features with Multiple Space Learning for Text-Based Video Retrieval

Damianos Galanopoulos, Vasileios Mezaris

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted for publication; to be included in Proc. ECCV Workshops 2022. The version posted here is the "submitted manuscript" version

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06119 2022-11-18 cs.CV 57%

SSGVS: Semantic Scene Graph-to-Video Synthesis

Yuren Cong, Jinhui Yi, Bodo Rosenhahn, Michael Ying Yang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08903 2022-11-17 cs.LG cs.AI 57%

Cross-Mode Knowledge Adaptation for Bike Sharing Demand Prediction using Domain-Adversarial Graph Neural Networks

Yuebing Liang, Guan Huang, Zhan Zhao

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2203.10961

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07216 2022-11-11 cs.CV cs.LG 57%

Class-attention Video Transformer for Engagement Intensity Prediction

Xusheng Ai, Victor S. Sheng, Chunhua Li, Zhiming Cui

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.08109 2022-11-03 cs.CV 57%

Boundary Proposal Network for Two-Stage Natural Language Video Localization

Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, Jun Xiao

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00509 2022-11-02 cs.CV 57%

Self-Supervised Intensity-Event Stereo Matching

Jinjin Gu, Jinan Zhou, Ringo Sai Wo Chu, Yan Chen, Jiawei Zhang, Xuanye Cheng, Song Zhang, Jimmy S. Ren

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments This paper has been accepted by the Journal of Imaging Science & Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12243 2022-11-02 cs.CV 57%

Safety-compliant Generative Adversarial Networks for Human Trajectory Forecasting

Parth Kothari, Alexandre Alahi

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 12 pages, 7 figures, 8 tables; Added acknowledgement

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16059 2022-10-31 cs.AI cs.CY 57%

An Artificial Intelligence driven Learning Analytics Method to Examine the Collaborative Problem solving Process from a Complex Adaptive Systems Perspective

Fan Ouyang, Weiqi Xu, Mutlu Cukurova

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 27 pages, 8 Figures, Accepted to appear in the International Journal of Computer-Supported Collaborative Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12617 2022-10-25 cs.CL 57%

Modal-specific Pseudo Query Generation for Video Corpus Moment Retrieval

Minjoon Jung, Seongho Choi, Joochan Kim, Jin-Hwa Kim, Byoung-Tak Zhang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted by EMNLP 2022 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11933 2022-10-24 cs.CV 57%

Fine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding

Yuechen Wang, Wengang Zhou, Houqiang Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 11 pages, 4 figures, accepted by Findings of EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.07131 2022-10-24 cs.CV 57%

Leveraging Real Talking Faces via Self-Supervision for Robust Forgery Detection

Alexandros Haliassos, Rodrigo Mira, Stavros Petridis, Maja Pantic

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments CVPR 2022. Code: https://github.com/ahaliassos/RealForensics

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09441 2022-10-19 cs.CV cs.HC cs.RO 57%

Real-Time Driver Monitoring Systems through Modality and View Analysis

Yiming Ma, Victor Sanchez, Soodeh Nikan, Devesh Upadhyay, Bhushan Atote, Tanaya Guha

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Paper summaries that our work on the DAD dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06192 2022-10-13 cs.CV eess.IV eess.SP 57%

Pose-Guided Graph Convolutional Networks for Skeleton-Based Action Recognition

Han Chen, Yifan Jiang, Hanseok Ko

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05039 2022-10-12 cs.LG cs.CV 57%

Contrastive Video-Language Learning with Fine-grained Frame Sampling

Zixu Wang, Yujie Zhong, Yishu Miao, Lin Ma, Lucia Specia

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments AACL-IJCNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04829 2022-10-11 cs.CL 57%

Hierarchical3D Adapters for Long Video-to-text Summarization

Pinelopi Papalampidi, Mirella Lapata

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04801 2022-10-11 cs.CV 57%

4D Unsupervised Object Discovery

Yuqi Wang, Yuntao Chen, Zhaoxiang Zhang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2022. 17 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.15050 2022-10-11 cs.CV 57%

AssistSR: Task-oriented Video Segment Retrieval for Personal AI Assistant

Stan Weixian Lei, Difei Gao, Yuxuan Wang, Dongxing Mao, Zihan Liang, Lingmin Ran, Mike Zheng Shou

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 20 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏