arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

1910.08593 2020-03-30 eess.IV cs.CV 57%

Generative Adversarial Networks And Domain Adaptation For Training Data Independent Image Registration

Dwarikanath Mahapatra

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.06891 2020-03-26 cs.CV 57%

Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences

Zhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang, Huasheng Liu, Lianli Gao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments The camera ready version for CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.07623 2020-03-18 cs.CV cs.LG 57%

Anomaly Detection in Video Data Based on Probabilistic Latent Space Models

Giulia Slavic, Damian Campo, Mohamad Baydoun, Pablo Marin, David Martin, Lucio Marcenaro, Carlo Regazzoni

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.12177 2020-02-28 cs.CV cs.LG 57%

Evolving Losses for Unsupervised Video Representation Learning

AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:1906.03248

Journal ref CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.08097 2020-02-20 cs.CV 57%

Unsupervised Temporal Feature Aggregation for Event Detection in Unstructured Sports Videos

Subhajit Chaudhury, Daiki Kimura, Phongtharin Vinayavekhin, Asim Munawar, Ryuki Tachibana, Koji Ito, Yuki Inaba, Minoru Matsumoto, Shuji Kidokoro, Hiroki Ozaki

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to IEEE International Symposium on Multimedia, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.10751 2020-02-19 cs.CV 57%

Deep Image-to-Video Adaptation and Fusion Networks for Action Recognition

Yang Liu, Zhaoyang Lu, Jing Li, Tao Yang, Chao Yao

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing, codes can be found at https://yangliu9208.github.io/DIVAFN/

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.11515 2020-02-19 cs.CV 57%

RhythmNet: End-to-end Heart Rate Estimation from Face via Spatial-temporal Representation

Xuesong Niu, Shiguang Shan, Hu Han, Xilin Chen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.13487 2020-02-17 cs.CV 57%

Use What You Have: Video Retrieval Using Representations From Collaborative Experts

Yang Liu, Samuel Albanie, Arsha Nagrani, Andrew Zisserman

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments This update contains a correction to previously reported results

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.05979 2020-01-17 cs.CV 57%

Contextual Sense Making by Fusing Scene Classification, Detections, and Events in Full Motion Video

Marc Bosch, Joseph Nassar, Benjamin Ortiz, Brendan Lammers, David Lindenbaum, John Wahl, Robert Mangum, Margaret Smith

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.05833 2020-01-17 cs.CV eess.IV 57%

Short-Term Temporal Convolutional Networks for Dynamic Hand Gesture Recognition

Yi Zhang, Chong Wang, Ye Zheng, Jieyu Zhao, Yuqi Li, Xijiong Xie

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.08936 2019-12-20 cs.CV 57%

One-Shot Weakly Supervised Video Object Segmentation

Mennatullah Siam, Naren Doraiswamy, Boris N. Oreshkin, Hengshuai Yao, Martin Jagersand

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.12283 2019-10-29 cs.AI 57%

Long-term Joint Scheduling for Urban Traffic

Xianfeng Liang, Likang Wu, Joya Chen, Yang Liu, Runlong Yu, Min Hou, Han Wu, Yuyang Ye, Qi Liu, Enhong Chen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments KDD Cup 2019 Special PaddlePaddle Award

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.13162 2019-10-01 cs.CL cs.LG 57%

Translation, Sentiment and Voices: A Computational Model to Translate and Analyse Voices from Real-Time Video Calling

Aneek Barman Roy

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CL

Comments 79 Pages, 19 Tables, 24 Figures, A M.Sc Dissertation

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.12948 2019-10-01 cs.CV eess.IV 57%

Video Skimming: Taxonomy and Comprehensive Survey

Vivekraj V. K., Debashis Sen, Balasubramanian Raman

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Journal ref ACM Computing Surveys (CSUR), Volume 52, Issue 5, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.05743 2019-10-01 cs.LG cs.CV stat.ML 57%

Learning Video Representations using Contrastive Bidirectional Transformer

Chen Sun, Fabien Baradel, Kevin Murphy, Cordelia Schmid

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.09283 2019-09-23 cs.CV 57%

Coupled Generative Adversarial Network for Continuous Fine-grained Action Segmentation

Harshala Gammulle, Tharindu Fernando, Simon Denman, Sridha Sridharan, Clinton Fookes

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments WACV 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.07137 2019-09-17 cs.CV 57%

PLIN: A Network for Pseudo-LiDAR Point Cloud Interpolation

Haojie Liu, Kang Liao, Chunyu Lin, Yao Zhao, Yulan Guo

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 7 pages, 5 figures, Submitted to ICRA2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.09540 2019-08-30 cs.CV 57%

Uncertainty-Aware Anticipation of Activities

Yazan Abu Farha, Juergen Gall

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments International Workshop on Human Behaviour Understanding, in conjunction with ICCV 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.00579 2019-08-09 cs.MM cs.IR 57%

Herding Effect based Attention for Personalized Time-Sync Video Recommendation

Wenmian Yang, Wenyuan Gao, Xiaojie Zhou, Weijia Jia, Shaohua Zhang, Yutao Luo

专题命中 视频多模态 :image-text(abstract);分类 cs.MM

Comments ACCEPTED for ORAL presentation at IEEE ICME 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.01311 2019-08-06 cs.CV 57%

Fully Automatic Video Colorization with Self-Regularization and Diversity

Chenyang Lei, Qifeng Chen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Published at the Computer Vision and Pattern Recognition (CVPR), 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.00744 2019-08-05 cs.HC cs.AI cs.RO 57%

Towards Learning How to Properly Play UNO with the iCub Robot

Pablo Barros, Stefan Wermter, Alessandra Sciutti

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Workshops on Naturalistic Non-Verbal and Affective Human-Robot Interactions co-located with ICDL-EPIROB 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.13369 2019-08-05 cs.CV 57%

Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video Recognition

Wenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen, Shilei Wen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2019 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.11117 2019-08-02 cs.CV 57%

Learning Visual Actions Using Multiple Verb-Only Labels

Michael Wray, Dima Damen

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at BMVC 2019. More information can be found at https://mwray.github.io/MVOL/. Annotations can be found at https://github.com/mwray/Multi-Verb-Labels

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.08437 2019-07-29 cs.CV 57%

Learning with privileged information via adversarial discriminative modality distillation

Nuno C. Garcia, Pietro Morerio, Vittorio Murino

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.11062 2019-07-26 cs.CL 57%

HireNet: a Hierarchical Attention Model for the Automatic Analysis of Asynchronous Video Job Interviews

Léo Hemamou, Ghazi Felhi, Vincent Vandenbussche, Jean-Claude Martin, Chloé Clavel

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments AAAI 2019

Journal ref Vol 33 (2019): Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.12158 2019-07-01 cs.CV cs.LG 57%

Open-Ended Long-Form Video Question Answering via Hierarchical Convolutional Self-Attention Networks

Zhu Zhang, Zhou Zhao, Zhijie Lin, Jingkuan Song, Xiaofei He

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by IJCAI 2019 as a poster paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.08861 2019-06-24 cs.NE cs.CV cs.LG eess.IV stat.ML 57%

Synthesizing Images from Spatio-Temporal Representations using Spike-based Backpropagation

Deboleena Roy, Priyadarshini Panda, Kaushik Roy

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 17 pages, 10 Figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.03248 2019-06-10 cs.CV 57%

Evolving Losses for Unlabeled Video Representation Learning

AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Non-archival abstract for CVPR Workshop on Learning from Unlabeled Videos

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.03782 2019-06-10 cs.CV 57%

Long-Term Occupancy Grid Prediction Using Recurrent Neural Networks

Marcel Schreiber, Stefan Hoermann, Klaus Dietmayer

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 8 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.04297 2019-04-10 cs.CV 57%

Learned 3D Shape Representations Using Fused Geometrically Augmented Images: Application to Facial Expression and Action Unit Detection

Bilal Taha, Munawar Hayat, Stefano Berretti, Naoufel Werghi

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏