arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4559 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4559 篇

2012.01002 2020-12-03 cs.CL cs.LG 79%

Classification of Multimodal Hate Speech -- The Winning Solution of Hateful Memes Challenge

Xiayu Zhong

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.06087 2020-11-13 cs.HC cs.MM 79%

Evoking Places from Spaces. The application of multimodal narrative techniques in the creation of "U Modified"

Gareth W. Young, Siobhán Mannion, Sara Wentworth

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments 5 pages

Journal ref 15th Sound and Music Computing Conference (SMC2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.02132 2020-11-05 eess.AS 79%

Multi-Modal Transformers Utterance-Level Code-Switching Detection

Krishna D N

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.01819 2020-11-04 cs.CV 79%

Learning Representations from Audio-Visual Spatial Alignment

Pedro Morgado, Yi Li, Nuno Vasconcelos

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments To appear at Advances in Neural Information Processing Systems (NeurIPS), 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.05019 2020-11-03 cs.CL cs.LG 79%

Multi-modal embeddings using multi-task learning for emotion recognition

Aparna Khare, Srinivas Parthasarathy, Shiva Sundaram

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments To appear in Interspeech,2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.08642 2020-10-20 cs.CL 79%

Multimodal Speech Recognition with Unstructured Audio Masking

Tejas Srinivasan, Ramon Sanabria, Florian Metze, Desmond Elliott

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to NLP Beyond Text workshop, EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.06318 2020-10-14 cs.CV 79%

Audio-Visual Self-Supervised Terrain Type Discovery for Mobile Platforms

Akiyoshi Kurobe, Yoshikatsu Nakajima, Hideo Saito, Kris Kitani

专题命中 音频语音多模态 :audio-visual(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.04470 2020-10-12 cs.CV 79%

gundapusunil at SemEval-2020 Task 8: Multimodal Memotion Analysis

Sunil Gundapu, Radhika Mamidi

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02384 2020-10-07 cs.CL 79%

Fine-Grained Grounding for Multimodal Speech Recognition

Tejas Srinivasan, Ramon Sanabria, Florian Metze, Desmond Elliott

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to Findings of EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.05651 2020-10-07 cs.AI cs.LG 79%

Multimodal Depression Severity Prediction from medical bio-markers using Machine Learning Tools and Technologies

Shivani Shimpi, Shyam Thombre, Snehal Reddy, Ritik Sharma, Srijan Singh

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments The paper content has one major problem in a section that is needs more attention and some work, and is agreed by all authors. The change is necessary in terms of correctness and completeness of the paper. Hence, requesting for the withdrawal to allow for the changes

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.14198 2020-10-06 cs.CL 79%

Multimodal Routing: Improving Local and Global Interpretability of Multimodal Language Analysis

Yao-Hung Hubert Tsai, Martin Q. Ma, Muqiao Yang, Ruslan Salakhutdinov, Louis-Philippe Morency

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.14361 2020-10-01 cs.CL cs.CY 79%

Ethically Collecting Multi-Modal Spontaneous Conversations with People that have Cognitive Impairments

Angus Addlesee, Pierre Albert

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments Published at LREC's Workshop on Legal and Ethical Issues in Human Language Technologies 2020

Journal ref LREC Workshop on Legal and Ethical Issues in Human Language Technologies (2020) 15-20

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.00402 2020-09-02 cs.CV 79%

Multimodal Aggregation Approach for Memory Vision-Voice Indoor Navigation with Meta-Learning

Liqi Yan, Dongfang Liu, Yaoxian Song, Changbin Yu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures, 2 tables, accepted at 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.07473 2020-08-07 cs.CV 79%

Dual-modality seq2seq network for audio-visual event localization

Yan-Bo Lin, Yu-Jhe Li, Yu-Chiang Frank Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted in ICASSP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.06809 2020-07-22 cs.CV 79%

DeepMSRF: A novel Deep Multimodal Speaker Recognition framework with Feature selection

Ehsan Asali, Farzan Shenavarmasouleh, Farid Ghareh Mohammadi, Prasanth Sengadu Suresh, Hamid R. Arabnia

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments The 24th International Conference on Image Processing, Computer Vision, & Pattern Recognition (IPCV'20: July 27-30, 2020, USA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.09903 2020-07-21 cs.CL 79%

Multimodal Dialogue State Tracking By QA Approach with Data Augmentation

Xiangyang Mou, Brandyn Sigouin, Ian Steenstra, Hui Su

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments AAAI DSTC8 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.14646 2020-06-01 eess.AS 79%

The INESC-ID Multi-Modal System for the ADReSS 2020 Challenge

Anna Pompili, Thomas Rolland, Alberto Abad

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments 5 pages, 1 figure. Submitted to INTERSPEECH2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.08335 2020-05-19 eess.AS cs.SD 79%

Multimodal Target Speech Separation with Voice and Face References

Leyuan Qu, Cornelius Weber, Stefan Wermter

专题命中 音频语音多模态 :multimodal(title);audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.06589 2020-05-14 cs.CV 79%

Arbitrary Talking Face Generation via Attentional Audio-Visual Coherence Learning

Hao Zhu, Huaibo Huang, Yi Li, Aihua Zheng, Ran He

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments IJCAI-2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.13236 2020-04-29 cs.CV 79%

Deep Auto-Encoders with Sequential Learning for Multimodal Dimensional Emotion Recognition

Dung Nguyen, Duc Thanh Nguyen, Rui Zeng, Thanh Thi Nguyen, Son N. Tran, Thin Nguyen, Sridha Sridharan, Clinton Fookes

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments Under Review on Transaction on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.01067 2020-04-15 cs.LG cs.SD eess.AS stat.ML 79%

Multimodal Deep Learning for Mental Disorders Prediction from Audio Speech Samples

Habibeh Naderi, Behrouz Haji Soleimani, Stan Matwin

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments arXiv admin note: text overlap with arXiv:1811.09362 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.12265 2020-03-30 cs.MM cs.IR cs.LG 79%

Unsupervised Cross-Modal Audio Representation Learning from Unstructured Multilingual Text

Alexander Schindler, Sergiu Gordea, Peter Knees

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.MM

Comments This is the long version of our SAC2020 poster presentation

Journal ref In Proceedings of the 35th ACM/SIGAPP Symposium On Applied Computing (SAC2020), March 30-April 3, 2020, Brno, Czech Republic

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.06656 2020-03-17 eess.AS cs.SD eess.IV 79%

Audio-Visual Spatial Aligment Requirements of Central and Peripheral Object Events

Davide Berghi, Hanne Stenzel, Marco Volino, Adrian Hilton, Philip J. B. Jackson

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Two-pages poster abstract

Journal ref IEEE VR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.12766 2020-03-02 eess.AS cs.SD 79%

Multi-Modal Continuous Valence And Arousal Prediction in the Wild Using Deep 3D Features and Sequence Modeling

Sowmya Rasipuram, Junaid Hamid Bhat, Anutosh Maitra

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.00854 2020-02-25 eess.AS cs.SD 79%

Re-synchronization using the Hand Preceding Model for Multi-modal Fusion in Automatic Continuous Cued Speech Recognition

Li Liu, Gang Feng, Denis Beautemps, Xiao-Ping Zhang

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Journal ref IEEE TMM-2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.04758 2020-01-15 cs.CV 79%

Deep Audio-Visual Learning: A Survey

Hao Zhu, Mandi Luo, Rui Wang, Aihua Zheng, Ran He

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.10693 2020-01-09 cs.CV 79%

DAVE: A Deep Audio-Visual Embedding for Dynamic Saliency Prediction

Hamed R. Tavakoli, Ali Borji, Esa Rahtu, Juho Kannala

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.01656 2020-01-07 eess.AS cs.SD 79%

Audio-visual Recognition of Overlapped speech for the LRS2 dataset

Jianwei Yu, Shi-Xiong Zhang, Jian Wu, Shahram Ghorbani, Bo Wu, Shiyin Kang, Shansong Liu, Xunying Liu, Helen Meng, Dong Yu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments 5 pages, 5 figures, submitted to icassp2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.12134 2019-12-30 cs.CV eess.SP 79%

Large-scale Multi-modal Person Identification in Real Unconstrained Environments

Jiajie Ye, Yisheng Guan, Junfa Liu, Xinghong Huang, Hong Zhang

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 6 pages, IEEE International Conference on Robotics and Biomimetics 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.12798 2019-12-02 cs.CL 79%

Multimodal Machine Translation through Visuals and Speech

Umut Sulubacak, Ozan Caglayan, Stig-Arne Grönroos, Aku Rouhe, Desmond Elliott, Lucia Specia, Jörg Tiedemann

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 34 pages, 4 tables, 8 figures. Submitted (Nov 2019) to the Machine Translation journal (Springer)

详情

展开后加载摘要…

URL PDF HTML 收藏