arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2106.14137 2021-06-29 cs.CV 79%

Building a Video-and-Language Dataset with Human Actions for Multimodal Logical Inference

Riko Suzuki, Hitomi Yanaka, Koji Mineshima, Daisuke Bekki

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to MMSR I

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08252 2021-05-19 cs.CV 79%

Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal Matching

Bofeng Wu, Guocheng Niu, Jun Yu, Xinyan Xiao, Jian Zhang, Hua Wu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.08833 2021-05-04 cs.CV 79%

Skeleton Aware Multi-modal Sign Language Recognition

Songyao Jiang, Bin Sun, Lichen Wang, Yue Bai, Kunpeng Li, Yun Fu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments This is a preprint version of our work SAM-SLR that ranked 1st at CVPR2021 Challenge on Large Scale Signer Independent Isolated Sign Language Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10139 2021-04-21 cs.CL 79%

Towards Solving Multimodal Comprehension

Pritish Sahu, Karan Sikka, Ajay Divakaran

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.03086 2021-04-06 eess.IV cs.CV 79%

A Multi-Modal Respiratory Disease Exacerbation Prediction Technique Based on a Spatio-Temporal Machine Learning Architecture

Rohan Tan Bhowmik

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Updated Title, References, and Acknowledgements

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.06747 2021-03-30 cs.CV 79%

ChallenCap: Monocular 3D Capture of Challenging Human Performances using Multi-Modal References

Yannan He, Anqi Pang, Xin Chen, Han Liang, Minye Wu, Yuexin Ma, Lan Xu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10572 2021-03-23 cs.MM 79%

Quantum-inspired Multimodal Fusion for Video Sentiment Analysis

Qiuchi Li, Dimitris Gkoumas, Christina Lioma, Massimo Melucci

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments Post-print accepted by Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.09320 2021-02-19 cs.CV 79%

Combining Events and Frames using Recurrent Asynchronous Multimodal Networks for Monocular Depth Prediction

Daniel Gehrig, Michelle Rüegg, Mathias Gehrig, Javier Hidalgo Carrio, Davide Scaramuzza

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref IEEE Robotics and Automation Letters (RA-L), 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.04762 2021-02-10 cs.CV 79%

Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

Linwei Ye, Mrigank Rochan, Zhi Liu, Xiaoqin Zhang, Yang Wang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 14 pages, 8 figures. arXiv admin note: substantial text overlap with arXiv:1904.04745

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.04727 2021-02-10 cs.CV 79%

Fashion Focus: Multi-modal Retrieval System for Video Commodity Localization in E-commerce

Yanhao Zhang, Qiang Wang, Pan Pan, Yun Zheng, Cheng Da, Siyang Sun, Yinghui Xu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments accepted by AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.08165 2021-01-21 cs.CV 79%

Video Relation Detection with Trajectory-aware Multi-modal Features

Wentao Xie, Guanghui Ren, Si Liu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.07339 2021-01-21 cs.CL 79%

MONAH: Multi-Modal Narratives for Humans to analyze conversations

Joshua Y. Kim, Greyson Y. Kim, Chunfeng Liu, Rafael A. Calvo, Silas C. R. Taylor, Kalina Yacef

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.11851 2020-12-23 cs.CV 79%

Predicting Online Video Advertising Effects with Multimodal Deep Learning

Jun Ikeda, Hiroyuki Seshime, Xueting Wang, Toshihiko Yamasaki

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at International Conference on Pattern Recognition 2020 (ICPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.03186 2020-12-11 cs.CV 79%

Noise Estimation Using Density Estimation for Self-Supervised Multimodal Learning

Elad Amrani, Rami Ben-Ari, Daniel Rotman, Alex Bronstein

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00514 2020-12-02 cs.CV cs.RO 79%

Multi-Modal Hybrid Architecture for Pedestrian Action Prediction

Amir Rasouli, Tiffany Yau, Mohsen Rohani, Jun Luo

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 7 pages, 4 Figures, 3 tables, submitted to ICRA 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.04417 2020-11-11 cs.CV 79%

Disentangle, align and fuse for multimodal and semi-supervised image segmentation

Agisilaos Chartsias, Giorgos Papanastasiou, Chengjia Wang, Scott Semple, David E. Newby, Rohan Dharmakumar, Sotirios A. Tsaftaris

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref IEEE Transactions on Medical Imaging (2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.13839 2020-10-28 cs.LG cs.CL 79%

VisualHints: A Visual-Lingual Environment for Multimodal Reinforcement Learning

Thomas Carta, Subhajit Chaudhury, Kartik Talamadupula, Michiaki Tatsubori

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Code is available at http://ibm.biz/VisualHints

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.13626 2020-10-27 cs.CV cs.LG 79%

Classification of Important Segments in Educational Videos using Multimodal Features

Junaid Ahmed Ghauri, Sherzod Hakimov, Ralph Ewerth

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Proceedings of the CIKM 2020 Workshops, October 19 to 20, Galway, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.11884 2020-10-23 cs.CV cs.HC 79%

AEGIS: A real-time multimodal augmented reality computer vision based system to assist facial expression recognition for individuals with autism spectrum disorder

James Ren Hou Lee, Alexander Wong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 4 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.08018 2020-09-18 cs.CL cs.IR 79%

Multi-modal Summarization for Video-containing Documents

Xiyan Fu, Jun Wang, Zhenglu Yang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.05702 2020-09-15 cs.RO cs.AI cs.LG cs.SY eess.SY 79%

Risk-Sensitive Sequential Action Control with Multi-Modal Human Trajectory Forecasting for Safe Crowd-Robot Interaction

Haruki Nishimura, Boris Ivanovic, Adrien Gaidon, Marco Pavone, Mac Schwager

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments To appear in 2020 IEEE/RSJ IROS

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.07935 2020-08-25 cs.CV 79%

Describing Unseen Videos via Multi-Modal Cooperative Dialog Agents

Ye Zhu, Yu Wu, Yi Yang, Yan Yan

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments ECCV2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.01148 2020-08-17 cs.RO cs.CV 79%

HAMLET: A Hierarchical Multimodal Attention-based Human Activity Recognition Algorithm

Md Mofijul Islam, Tariq Iqbal

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments To be published in the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2020

Journal ref IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02036 2020-07-07 cs.CV 79%

Modality Shifting Attention Network for Multi-modal Video Question Answering

Junyeong Kim, Minuk Ma, Trung Pham, Kyungsu Kim, Chang D. Yoo

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments CVPR2020 accepted; poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.03844 2020-06-09 cs.CV cs.LG stat.ML 79%

Exploiting Temporal Coherence for Multi-modal Video Categorization

Palash Goyal, Saurabh Sahu, Shalini Ghosh, Chul Lee

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.00654 2020-06-02 cs.LG cs.CV stat.ML 79%

A multimodal approach for multi-label movie genre classification

Rafael B. Mangolin, Rodolfo M. Pereira, Alceu S. Britto, Carlos N. Silla, Valéria D. Feltrim, Diego Bertolini, Yandre M. G. Costa

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 21 pages and 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.13876 2020-05-29 cs.MM 79%

Investigating Correlations of Automatically Extracted Multimodal Features and Lecture Video Quality

Jianwei Shi, Christian Otto, Anett Hoppe, Peter Holtz, Ralph Ewerth

专题命中 视频多模态 :multimodal(title);cross-modal(abstract);分类 cs.MM

Journal ref SALMM '19: Proceedings of the 1st International Workshop on Search as Learning with Multimedia Information, co-located with ACM Multimedia 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.09606 2020-05-20 cs.CL 79%

A Recipe for Creating Multimodal Aligned Datasets for Sequential Tasks

Angela S. Lin, Sudha Rao, Asli Celikyilmaz, Elnaz Nouri, Chris Brockett, Debadeepta Dey, Bill Dolan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments This paper has been accepted to be published at ACL 2020

Journal ref Association of Computational Linguistics 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.08637 2020-05-20 cs.HC cs.CV 79%

Building BROOK: A Multi-modal and Facial Video Database for Human-Vehicle Interaction Research

Xiangjun Peng, Zhentao Huang, Xu Sun

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Conference: ACM CHI Conference on Human Factors in Computing Systems Workshops (CHI'20 Workshops)At: Honolulu, Hawaii, USA URL:https://emergentdatatrails.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.06502 2020-04-15 cs.CV cs.LG eess.IV 79%

Unsupervised Multimodal Video-to-Video Translation via Self-Supervised Learning

Kangning Liu, Shuhang Gu, Andres Romero, Radu Timofte

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏