arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

共收录 495 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 音视频/视觉语言融合 495 篇

2106.11411 2021-06-23 cs.SD eess.AS 50%

Attention-based cross-modal fusion for audio-visual voice activity detection in musical video streams

Yuanbo Hou, Zhesong Yu, Xia Liang, Xingjian Du, Bilei Zhu, Zejun Ma, Dick Botteldooren

专题命中 音视频/视觉语言融合 :information fusion(abstract)

Comments Accepted by INTERSPEECH 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.11059 2021-06-22 cs.LG 50%

Improving Multi-Modal Learning with Uni-Modal Teachers

Chenzhuang Du, Tingle Li, Yichen Liu, Zixin Wen, Tianyu Hua, Yue Wang, Hang Zhao

专题命中 音视频/视觉语言融合 :multi-modal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10121 2021-04-21 cs.SD cs.CL eess.AS 50%

On the Impact of Word Error Rate on Acoustic-Linguistic Speech Emotion Recognition: An Update for the Deep Learning Era

Shahin Amiriparian, Artem Sokolov, Ilhan Aslan, Lukas Christ, Maurice Gerczuk, Tobias Hübner, Dmitry Lamanov, Manuel Milling, Sandra Ottl, Ilya Poduremennykh, Evgeniy Shuranov, Björn W. Schuller

专题命中 音视频/视觉语言融合 :decision-level fusion(abstract)

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.01894 2021-03-03 cs.SD cs.CL eess.AS 50%

Investigations on Audiovisual Emotion Recognition in Noisy Conditions

Michael Neumann, Ngoc Thang Vu

专题命中 音视频/视觉语言融合 :hybrid fusion(abstract)

Comments Published at the IEEE workshop on Spoken Language Technology (SLT) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.01326 2021-02-03 eess.AS cs.LG cs.SD 50%

Multimodal Attention Fusion for Target Speaker Extraction

Hiroshi Sato, Tsubasa Ochiai, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani, Shoko Araki

专题命中 音视频/视觉语言融合 :multi-modal fusion(abstract)

Comments 7 pages, 5 figures

Journal ref in IEEE Spoken Language Technology Workshop (SLT), 2021, pp. 778-784

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.13449 2020-12-29 cs.HC 50%

You Have a Point There: Object Selection Inside an Automobile Using Gaze, Head Pose and Finger Pointing

Abdul Rafey Aftab, Michael von der Beeck, Michael Feld

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Journal ref In Proceedings of the 2020 International Conference on Multimodal Interaction, pp. 595-603. 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.06584 2020-12-21 cs.HC 50%

Jointly Optimizing Sensing Pipelines for Multimodal Mixed Reality Interaction

Darshana Rathnayake, Ashen de Silva, Dasun Puwakdandawa, Lakmal Meegahapola, Archan Misra, Indika Perera

专题命中 音视频/视觉语言融合 :sensor fusion(abstract)

Comments 17th IEEE International Conference on Mobile Ad-Hoc and Sensor Systems (MASS) - 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.00702 2020-08-04 eess.AS cs.CL 50%

Multimodal Semi-supervised Learning Framework for Punctuation Prediction in Conversational Speech

Monica Sunkara, Srikanth Ronanki, Dhanush Bekal, Sravan Bodapati, Katrin Kirchhoff

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments Accepted for Interspeech 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.12788 2020-07-21 eess.AS cs.AI 50%

Identification of Dementia Using Audio Biomarkers

Rupayan Chakraborty, Meghna Pandharipande, Chitralekha Bhat, Sunil Kumar Kopparapu

专题命中 音视频/视觉语言融合 :feature-level fusion(abstract)

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.08250 2020-04-20 eess.AS cs.LG 50%

How to Teach DNNs to Pay Attention to the Visual Modality in Speech Recognition

George Sterpu, Christian Saam, Naomi Harte

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments in IEEE/ACM Transactions on Audio, Speech, and Language Processing (to appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.01067 2020-04-15 cs.LG cs.SD eess.AS stat.ML 50%

Multimodal Deep Learning for Mental Disorders Prediction from Audio Speech Samples

Habibeh Naderi, Behrouz Haji Soleimani, Stan Matwin

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments arXiv admin note: text overlap with arXiv:1811.09362 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.10671 2019-10-24 cs.CL cs.LG eess.AS 50%

A practical two-stage training strategy for multi-stream end-to-end speech recognition

Ruizhi Li, Gregory Sell, Xiaofei Wang, Shinji Watanabe, Hynek Hermansky

专题命中 音视频/视觉语言融合 :information fusion(abstract)

Comments submitted to ICASSP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.11062 2019-07-26 cs.CL 50%

HireNet: a Hierarchical Attention Model for the Automatic Analysis of Asynchronous Video Job Interviews

Léo Hemamou, Ghazi Felhi, Vincent Vandenbussche, Jean-Claude Martin, Chloé Clavel

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments AAAI 2019

Journal ref Vol 33 (2019): Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.02844 2019-03-08 cs.SD eess.AS 50%

Voice Activity Detection: Merging Source and Filter-based Information

Thomas Drugman, Yannis Stylianou, Yusuke Kida, Masami Akamine

专题命中 音视频/视觉语言融合 :information fusion(abstract)

Journal ref IEEE Signal Processing Letters, Volume 23, Issue 2, pp. 252-256, 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.00924 2018-02-06 cs.LG cs.AI cs.CL stat.ML 50%

Multimodal Sentiment Analysis with Word-Level Fusion and Reinforcement Learning

Minghai Chen, Sen Wang, Paul Pu Liang, Tadas Baltrušaitis, Amir Zadeh, Louis-Philippe Morency

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments ICMI 2017 Oral Presentation, Honorable Mention Award

详情

展开后加载摘要…

URL PDF HTML 收藏