arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4557 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4557 篇

2303.02344 2023-03-07 cs.CV cs.MM 81%

Improving Audio-Visual Video Parsing with Pseudo Visual Labels

Jinxing Zhou, Dan Guo, Yiran Zhong, Meng Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00502 2023-03-02 cs.SD cs.CV eess.AS 81%

On the Audio-visual Synchronization for Lip-to-Speech Synthesis

Zhe Niu, Brian Mak

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13838 2023-03-02 cs.CV cs.SD eess.AS 81%

Cross-modal Face- and Voice-style Transfer

Naoya Takahashi, Mayank K. Singh, Yuki Mitsufuji

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07711 2023-03-02 cs.CL cs.AI cs.LG 81%

Multilevel Transformer For Multimodal Emotion Recognition

Junyi He, Meimei Wu, Meng Li, Xiaobo Zhu, Feng Ye

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 5 pages, 2 figures, 3 tables. Submitted to ICASSP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10315 2023-02-22 cs.CL cs.AI 81%

Generalization algorithm of multimodal pre-training model based on graph-text self-supervised training

Zhangxiaobing, Tangzhenhao, Longzi, Fuxianghua

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.05199 2023-02-14 cs.CL cs.MM cs.NE 81%

A Survey on Multi-modal Summarization

Anubhav Jangra, Sourajit Mukherjee, Adam Jatowt, Sriparna Saha, Mohammad Hasanuzzaman

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted in ACM CSUR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04156 2023-02-09 cs.CL cs.IR cs.MM 81%

Prompting for Multimodal Hateful Meme Classification

Rui Cao, Roy Ka-Wei Lee, Wen-Haw Chong, Jing Jiang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted in EMNLP, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.01555 2023-02-06 cs.AI cs.CV 81%

Bridging the Emotional Semantic Gap via Multimodal Relevance Estimation

Chuan Zhang, Daoxin Zhang, Ruixiu Zhang, Jiawei Li, Jianke Zhu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12202 2023-01-24 cs.SD cs.CV eess.AS physics.app-ph 81%

Multimodal Exponentially Modified Gaussian Oscillators

Christopher Hahne

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、eess.AS

Comments IEEE International Ultrasonic Symposium 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11868 2023-01-10 cs.CV cs.SD eess.AS 81%

Interpretable Multimodal Emotion Recognition using Hybrid Fusion of Speech and Image Data

Puneet Kumar, Sarthak Malik, Balasubramanian Raman

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、eess.AS

Comments arXiv admin note: text overlap with arXiv:2208.11450

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14843 2023-01-04 cs.SD cs.CV cs.LG cs.RO eess.AS 81%

Catch Me If You Hear Me: Audio-Visual Navigation in Complex Unmapped Environments with Moving Sounds

Abdelrahman Younes, Daniel Honerkamp, Tim Welschehold, Abhinav Valada

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

Comments This paper has been accepted for publication at IEEE ROBOTICS AND AUTOMATION LETTERS

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.11345 2022-12-23 cs.RO cs.AI cs.CV 81%

Knowledge-driven Scene Priors for Semantic Audio-Visual Embodied Navigation

Gyan Tatiya, Jonathan Francis, Luca Bondi, Ingrid Navarro, Eric Nyberg, Jivko Sinapov, Jean Oh

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.AI

Comments 19 pages, 8 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10196 2022-12-22 cs.CL cs.AI 81%

Multimodal Hate Speech Detection from Bengali Memes and Texts

Md. Rezaul Karim, Sumon Kanti Dey, Tanhim Islam, Md. Shajalal, Bharathi Raja Chakravarthi

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: text overlap with arXiv:2107.00648 by other authors

Journal ref Pre-print for our paper at International Conference on Speech & Language Technology for Low-resource Languages (SPELLL'2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09090 2022-12-20 cs.SD cs.MM eess.AS 81%

Exploring Workplace Behaviors through Speaking Patterns using Large-scale Multimodal Wearable Recordings: A Study of Healthcare Providers

Tiantian Feng, Shrikanth Narayanan

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08661 2022-12-20 cs.LG cs.AI cs.CL 81%

EffMulti: Efficiently Modeling Complex Multimodal Interactions for Emotion Analysis

Feng Qiu, Chengyang Xie, Yu Ding, Wanzeng Kong

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 6 pages,1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04970 2022-12-12 cs.CV cs.AI cs.GR 81%

Masked Lip-Sync Prediction by Audio-Visual Contextual Exploitation in Transformers

Yasheng Sun, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Zhibin Hong, Jingtuo Liu, Errui Ding, Jingdong Wang, Ziwei Liu, Hideki Koike

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to SIGGRAPH Asia 2022 (Conference Proceedings). Project page: https://hangz-nju-cuhk.github.io/projects/AV-CAT

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01892 2022-12-06 eess.AS cs.MM cs.SD 81%

Tragic Talkers: A Shakespearean Sound- and Light-Field Dataset for Audio-Visual Machine Learning Research

Davide Berghi, Marco Volino, Philip J. B. Jackson

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14700 2022-11-29 cs.CL eess.AS 81%

A novel multimodal dynamic fusion network for disfluency detection in spoken utterances

Sreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Manan Suri, Rajiv Ratn Shah

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments Submitted to ICASSP 2023. arXiv admin note: text overlap with arXiv:2203.16794

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04006 2022-11-28 cs.SD cs.CV cs.LG eess.AS 81%

Few-Shot Audio-Visual Learning of Environment Acoustics

Sagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen Grauman

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

Comments Accepted to NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08180 2022-11-23 cs.CL cs.LG cs.SD eess.AS 81%

SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation

Sameer Khurana, Antoine Laurent, James Glass

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03587 2022-11-08 cs.CV cs.AI cs.LG 81%

Generalized Product-of-Experts for Learning Multimodal Representations in Noisy Environments

Abhinav Joshi, Naman Gupta, Jinang Shah, Binod Bhattarai, Ashutosh Modi, Danail Stoyanov

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 11 Pages, Accepted at ICMI 2022 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00325 2022-11-02 eess.AS cs.CL cs.SD 81%

Speech-text based multi-modal training with bidirectional attention for improved speech recognition

Yuhang Yang, Haihua Xu, Hao Huang, Eng Siong Chng, Sheng Li

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL、eess.AS

Comments 5 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10439 2022-11-02 cs.CV cs.LG cs.SD eess.AS 81%

Transformer-Based Video Front-Ends for Audio-Visual Speech Recognition for Single and Multi-Person Video

Dmitriy Serdyuk, Otavio Braga, Olivier Siohan

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

Comments 5 pages, 3 figures, published at Interspeech 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16472 2022-11-01 cs.SD cs.CV cs.LG eess.AS 81%

Learning Audio-Visual Dynamics Using Scene Graphs for Audio Source Separation

Moitreya Chatterjee, Narendra Ahuja, Anoop Cherian

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

Comments Accepted at NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09556 2022-10-19 cs.CL cs.SD eess.AS 81%

Discrete Cross-Modal Alignment Enables Zero-Shot Speech Translation

Chen Wang, Yuchen Liu, Boxing Chen, Jiajun Zhang, Wei Luo, Zhongqiang Huang, Chengqing Zong

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CL、eess.AS

Comments Accepted by the main conference of EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.01353 2022-10-06 cs.SD cs.AI eess.AS 81%

Pay Self-Attention to Audio-Visual Navigation

Yinfeng Yu, Lele Cao, Fuchun Sun, Xiaohong Liu, Liejun Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI、eess.AS

Comments Main paper (10 pages and 7 figures) and appendix (21 figures and 4 tables). Accepted for publication by BMVC 2022. For data and code, see https://yyf17.github.io/FSAAVN/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.09634 2022-09-21 cs.SD cs.CV cs.LG cs.MM 81%

A Closer Look at Weakly-Supervised Audio-Visual Source Localization

Shentong Mo, Pedro Morgado

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10838 2022-08-24 cs.CV cs.CL 81%

Multimodal Crop Type Classification Fusing Multi-Spectral Satellite Time Series with Farmers Crop Rotations and Local Crop Distribution

Valentin Barriere, Martin Claverie

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments accepted to CECEO22@IJCAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.13354 2022-08-16 cs.MM cs.LG cs.SD eess.AS 81%

Audio-to-Image Cross-Modal Generation

Maciej Żelaszczyk, Jacek Mańdziuk

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.MM、eess.AS

Journal ref International Joint Conference on Neural Networks, IJCNN 2022, Padua, Italy, 1-8

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.03633 2022-08-09 cs.MM cs.SD eess.AS 81%

Debiased Cross-modal Matching for Content-based Micro-video Background Music Recommendation

Jinng Yi, Zhenzhong Chen

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏