arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2006.07364 2020-10-23 cs.RO cs.CV cs.GR cs.LG stat.ML 57%

Residual Force Control for Agile Human Behavior Imitation and Extended Motion Synthesis

Ye Yuan, Kris Kitani

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments NeurIPS 2020. Code: https://github.com/Khrylx/RFC. Project page: https://www.ye-yuan.com/rfc

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.05264 2020-10-13 cs.CV 57%

Boosting Continuous Sign Language Recognition via Cross Modality Augmentation

Junfu Pu, Wengang Zhou, Hezhen Hu, Houqiang Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to ACM Multimedia 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.09989 2020-10-13 cs.CV 57%

Characterizing the impact of using features extracted from pre-trained models on the quality of video captioning sequence-to-sequence models

Menatallh Hammad, May Hammad, Mohamed Elshenawy

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Submitted to conference ICPRAI2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.09878 2020-09-22 cs.CV cs.LG 57%

Haar Wavelet based Block Autoregressive Flows for Trajectories

Apratim Bhattacharyya, Christoph-Nikolas Straehle, Mario Fritz, Bernt Schiele

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments German Conference on Pattern Recognition, 2020 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.08614 2020-09-21 cs.CV 57%

Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos

Jie Wu, Guanbin Li, Xiaoguang Han, Liang Lin

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.07203 2020-09-21 cs.CV 57%

Video Understanding as Machine Translation

Bruno Korbar, Fabio Petroni, Rohit Girdhar, Lorenzo Torresani

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments The authors have temporarily withdrawn this paper to reassess some of the experimental results

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.08791 2020-08-21 cs.HC cs.CV eess.IV 57%

Facial movement synergies and Action Unit detection from distal wearable Electromyography and Computer Vision

Monica Perusquia-Hernandez, Felix Dollack, Chun Kwang Tan, Shushi Namba, Saho Ayabe-Kanamura, Kenji Suzuki

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 11 pages, 11 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.01403 2020-08-14 cs.CV cs.IR 57%

Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization

Daizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong, Pan Zhou, Zichuan Xu

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.02448 2020-08-07 cs.CV 57%

Fine-grained Iterative Attention Network for TemporalLanguage Localization in Videos

Xiaoye Qu, Pengwei Tang, Zhikang Zhou, Yu Cheng, Jianfeng Dong, Pan Zhou

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments ACM MM 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.15244 2020-07-31 cs.CV cs.LG 57%

Hierarchical Action Classification with Network Pruning

Mahdi Davoodikakhki, KangKang Yin

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.11346 2020-07-23 cs.HC cs.LG cs.MM 57%

Towards Social & Engaging Peer Learning: Predicting Backchanneling and Disengagement in Children

Mononito Goswami, Minkush Manuja, Maitree Leekha

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

Comments 14 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.10937 2020-07-22 cs.CV 57%

MovieNet: A Holistic Dataset for Movie Understanding

Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang, Dahua Lin

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ECCV2020 as spotlight presentation. Project page: http://movienet.site

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.01883 2020-07-07 cs.CV 57%

Egocentric Action Recognition by Video Attention and Temporal Context

Juan-Manuel Perez-Rua, Antoine Toisoul, Brais Martinez, Victor Escorcia, Li Zhang, Xiatian Zhu, Tao Xiang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments EPIC-Kitchens challenges@CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.01738 2020-07-06 cs.CV 57%

Video Prediction via Example Guidance

Jingwei Xu, Huazhe Xu, Bingbing Ni, Xiaokang Yang, Trevor Darrell

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Project Page: https://sites.google.com/view/vpeg-supp/home

Journal ref ICML 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.12799 2020-06-24 cs.CL 57%

Keyframe Segmentation and Positional Encoding for Video-guided Machine Translation Challenge 2020

Tosho Hirasawa, Zhishen Yang, Mamoru Komachi, Naoaki Okazaki

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments 4 pages; First Workshop on Advances in Language and Vision Research (ALVR 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.11393 2020-06-23 cs.CV cs.LG 57%

Unifying Few- and Zero-Shot Egocentric Action Recognition

Tyler R. Scott, Michael Shvartsman, Karl Ridgeway

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted for presentation at the EPIC@CVPR2020 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.08349 2020-06-23 q-bio.NC cs.CV eess.IV 57%

Investigating naturalistic hand movements by behavior mining in long-term video and neural recordings

Satpreet H. Singh, Steven M. Peterson, Rajesh P. N. Rao, Bingni W. Brunton

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.08505 2020-06-19 cs.CV cs.HC cs.LG cs.RO 57%

Dynamic Gesture Recognition by Using CNNs and Star RGB: a Temporal Information Condensation

Clebeson Canuto dos Santos, Jorge Leonid Aching Samatelo, Raquel Frizera Vassallo

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 19 pages, 12 figures, submitted to Neurocomputing Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.08748 2020-06-17 cs.CL 57%

DynE: Dynamic Ensemble Decoding for Multi-Document Summarization

Chris Hokamp, Demian Gholipour Ghalandari, Nghia The Pham, John Glover

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.00592 2020-06-11 cs.CY cs.AI cs.HC 57%

Predicting Engagement in Video Lectures

Sahan Bulathwela, María Pérez-Ortiz, Aldo Lipani, Emine Yilmaz, John Shawe-Taylor

专题命中 视频多模态 :cross-modal(abstract);分类 cs.AI

Comments In Proceedings of International Conference on Educational Data Mining 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.04859 2020-06-11 cs.CV cs.RO 57%

Novel Perception Algorithmic Framework For Object Identification and Tracking In Autonomous Navigation

Suryansh Saxena, Isaac K Isukapati

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.09540 2020-05-19 cs.CV 57%

Weakly Supervised Few-shot Object Segmentation using Co-Attention with Visual and Semantic Embeddings

Mennatullah Siam, Naren Doraiswamy, Boris N. Oreshkin, Hengshuai Yao, Martin Jagersand

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to IJCAI'20. The first three authors listed contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.02190 2020-05-11 cs.CV 57%

Rolling-Unrolling LSTMs for Action Anticipation from First-Person Video

Antonino Furnari, Giovanni Maria Farinella

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:1905.09035

Journal ref Published in IEEE Transaction on Pattern Analysis and Machine Interaction, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.13979 2020-04-30 cs.CV cs.LG eess.IV 57%

Skeleton Focused Human Activity Recognition in RGB Video

Bruce X. B. Yu, Yan Liu, Keith C. C. Chan

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.11994 2020-04-28 cs.MM cs.LG 57%

Bharatanatyam Dance Transcription using Multimedia Ontology and Machine Learning

Tanwi Mallick, Patha Pratim Das, Arun Kumar Majumdar

专题命中 视频多模态 :multi-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.03972 2020-04-21 eess.IV cs.CV cs.LG stat.ML 57%

IrisNet: Deep Learning for Automatic and Real-time Tongue Contour Tracking in Ultrasound Video Data using Peripheral Vision

M. Hamed Mozaffari, Md. Aminur Rab Ratul, Won-Sook Lee

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.05573 2020-04-14 cs.CV cs.LG 57%

YouMakeup VQA Challenge: Towards Fine-grained Action Understanding in Domain-Specific Videos

Shizhe Chen, Weiying Wang, Ludan Ruan, Linli Yao, Qin Jin

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments CVPR LVVU Workshop 2020 YouMakeup VQA Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.04959 2020-04-13 cs.MM cs.IR 57%

Stacked Convolutional Deep Encoding Network for Video-Text Retrieval

Rui Zhao, Kecheng Zheng, Zheng-jun Zha

专题命中 视频多模态 :cross-modal(abstract);分类 cs.MM

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.05878 2020-04-07 cs.IR cs.AI 57%

Event-Radar: Real-time Local Event Detection System for Geo-Tagged Tweet Streams

Sibo Zhang, Yuan Cheng, Deyuan Ke

专题命中 视频多模态 :cross-modal(abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.13158 2020-03-31 cs.CV 57%

Learning Interactions and Relationships between Movie Characters

Anna Kukleva, Makarand Tapaswi, Ivan Laptev

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR 2020 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏