arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2209.03609 2022-09-09 cs.CV cs.MM 81%

Frame-Subtitle Self-Supervision for Multi-Modal Video Question Answering

Jiong Wang, Zhou Zhao, Weike Jin

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.09079 2022-08-22 cs.LG cs.AI cs.CV cs.CY 81%

A Multi-Modal Wildfire Prediction and Personalized Early-Warning System Based on a Novel Machine Learning Framework

Rohan Tan Bhowmik

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.06773 2022-08-16 cs.CV cs.IR cs.LG cs.MM 81%

TL;DW? Summarizing Instructional Videos with Task Relevance & Cross-Modal Saliency

Medhini Narasimhan, Arsha Nagrani, Chen Sun, Michael Rubinstein, Trevor Darrell, Anna Rohrbach, Cordelia Schmid

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted to ECCV 2022. Website: https://medhini.github.io/ivsum/

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07077 2022-07-15 cs.CV cs.AI 81%

Egocentric Scene Understanding via Multimodal Spatial Rectifier

Tien Do, Khiem Vuong, Hyun Soo Park

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Appearing in the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02547 2022-04-07 cs.CV cs.CL 81%

Modeling Motion with Multi-Modal Features for Text-Based Video Segmentation

Wangbo Zhao, Kai Wang, Xiangxiang Chu, Fuzhao Xue, Xinchao Wang, Yang You

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14821 2022-04-05 cs.CV cs.CL cs.LG 81%

End-to-End Referring Video Object Segmentation with Multimodal Transformers

Adam Botach, Evgenii Zheltonozhskii, Chaim Baskin

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12441 2022-03-24 cs.AI cs.MM 81%

M-SENA: An Integrated Platform for Multimodal Sentiment Analysis

Huisheng Mao, Ziqi Yuan, Hua Xu, Wenmeng Yu, Yihe Liu, Kai Gao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments 11 pages, 4 figures, to be published in ACL 2022 System Demonstration Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02573 2022-03-08 cs.CV cs.AI cs.LG 81%

Show Me What and Tell Me How: Video Synthesis via Multimodal Conditioning

Ligong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri, Kyle Olszewski, Shervin Minaee, Dimitris Metaxas, Sergey Tulyakov

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.12724 2022-01-11 cs.LG cs.CL cs.CV 81%

Predicting the Popularity of Micro-videos with Multimodal Variational Encoder-Decoder Framework

Yaochen Zhu, Jiayi Xie, Zhenzhong Chen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03007 2022-01-10 cs.CL cs.LG cs.MM 81%

Unsupervised Multimodal Language Representations using Convolutional Autoencoders

Panagiotis Koromilas, Theodoros Giannakopoulos

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.12182 2021-12-24 cs.CV cs.AI 81%

Fine-grained Multi-Modal Self-Supervised Learning

Duo Wang, Salah Karout

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01849 2021-12-06 cs.MM cs.CV cs.LG 81%

Cross-modal Knowledge Distillation for Vision-to-Sensor Action Recognition

Jianyuan Ni, Raunak Sarbajna, Yang Liu, Anne H. H. Ngu, Yan Yan

专题命中 视频多模态 :cross-modal(title);multi-modal(abstract);分类 cs.CV、cs.MM

Comments 5 pages, 2 figures, submitted to ICASSP2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.01024 2021-11-02 cs.CV cs.SD eess.AS 81%

With a Little Help from my Temporal Context: Multimodal Egocentric Action Recognition

Evangelos Kazakos, Jaesung Huh, Arsha Nagrani, Andrew Zisserman, Dima Damen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、eess.AS

Comments Accepted at BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.02636 2021-10-25 cs.CV cs.CL cs.LG 81%

MERLOT: Multimodal Neural Script Knowledge Models

Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, Yejin Choi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments project page at https://rowanzellers.com/merlot; NeurIPS 2021 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09424 2021-10-19 cs.CV cs.CL cs.HC cs.LG 81%

Don't Judge Me by My Face : An Indirect Adversarial Approach to Remove Sensitive Information From Multimodal Neural Representation in Asynchronous Job Video Interviews

Léo Hemamou, Arthur Guillon, Jean-Claude Martin, Chloé Clavel

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments published in ACII 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06637 2021-09-15 cs.CV cs.MM 81%

Multi-modal Representation Learning for Video Advertisement Content Structuring

Daya Guo, Zhaoyang Zeng

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.08344 2021-08-20 cs.CV cs.AI 81%

The Multi-Modal Video Reasoning and Analyzing Competition

Haoran Peng, He Huang, Li Xu, Tianjiao Li, Jun Liu, Hossein Rahmani, Qiuhong Ke, Zhicheng Guo, Cong Wu, Rongchang Li, Mang Ye, Jiahao Wang, Jiaxu Zhang, Yuanzhong Liu, Tao He, Fuwei Zhang, Xianbin Liu, Tao Lin

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2021 Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.03354 2021-08-10 cs.LG cs.AI cs.HC cs.MM 81%

HetEmotionNet: Two-Stream Heterogeneous Graph Recurrent Neural Network for Multi-modal Emotion Recognition

Ziyu Jia, Youfang Lin, Jing Wang, Zhiyang Feng, Xiangheng Xie, Caijie Chen

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted by ACM MM 2021. The SOLE copyright holder is ACM Multimedia, all rights reserved

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.09525 2021-07-26 cs.CV cs.LG cs.SD eess.AS 81%

Toward Automated Classroom Observation: Multimodal Machine Learning to Estimate CLASS Positive Climate and Negative Climate

Anand Ramakrishnan, Brian Zylich, Erin Ottmar, Jennifer LoCasale-Crouch, Jacob Whitehill

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、eess.AS

Comments The authors discovered that the results are not reproducible

Journal ref IEEE Transactions on Affective Computing, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.05165 2021-05-13 cs.CV cs.AI cs.LG 81%

AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition

Rameswar Panda, Chun-Fu Chen, Quanfu Fan, Ximeng Sun, Kate Saenko, Aude Oliva, Rogerio Feris

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.10102 2021-04-30 cs.MM cs.CL 81%

A multimodal movie review corpus for fine-grained opinion mining

Alexandre Garcia, Slim Essid, Florence d'Alché-Buc, Chloé Clavel

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.09411 2021-04-20 cs.CV cs.MM 81%

Understanding Chinese Video and Language via Contrastive Multimodal Pre-Training

Chenyi Lei, Shixian Luo, Yong Liu, Wanggui He, Jiamang Wang, Guoxin Wang, Haihong Tang, Chunyan Miao, Houqiang Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.04853 2021-03-26 cs.CV cs.AI 81%

Social-STAGE: Spatio-Temporal Multi-Modal Future Trajectory Forecast

Srikanth Malla, Chiho Choi, Behzad Dariush

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments ICRA 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.11624 2021-03-25 cs.CV cs.AI 81%

Multimodal Motion Prediction with Stacked Transformers

Yicheng Liu, Jinghuai Zhang, Liangji Fang, Qinhong Jiang, Bolei Zhou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments CVPR2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.11704 2021-01-29 cs.LG cs.MM cs.SD eess.AS eess.IV 81%

A Case Study of Deep Learning Based Multi-Modal Methods for Predicting the Age-Suitability Rating of Movie Trailers

Mahsa Shafaei, Christos Smailis, Ioannis A. Kakadiaris, Thamar Solorio

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.00536 2021-01-06 cs.CV cs.AI 81%

A Multi-modal Machine Learning Approach and Toolkit to Automate Recognition of Early Stages of Dementia among British Sign Language Users

Xing Liang, Anastassia Angelopoulou, Epaminondas Kapetanios, Bencie Woll, Reda Al-batat, Tyron Woolfe

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Journal ref ECCV 2020 Workshops. Lecture Notes in Computer Science, Vol 12536. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.00073 2021-01-05 cs.CV cs.AI 81%

A Multi-modal Deep Learning Model for Video Thumbnail Selection

Zhifeng Yu, Nanchun Shi

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.10019 2021-01-05 cs.CV cs.AI 81%

Hierarchical Conditional Relation Networks for Multimodal Video Question Answering

Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Major extension of our CVPR'20 paper to handle long video with text. arXiv admin note: substantial text overlap with arXiv:2002.10698

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.09290 2021-01-01 cs.CV cs.MM 81%

Frame Aggregation and Multi-Modal Fusion Framework for Video-Based Person Recognition

Fangtao Li, Wenzhe Wang, Zihe Liu, Haoran Wang, Chenghao Yan, Bin Wu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by MMM 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.09046 2020-11-25 cs.CV cs.CL 81%

A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus

Bowen Zhang, Hexiang Hu, Joonseok Lee, Ming Zhao, Sheide Chammas, Vihan Jain, Eugene Ie, Fei Sha

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏