arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4559 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4559 篇

2312.14378 2024-02-12 cs.LG cs.SD eess.AS 79%

Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification

Anirudh S. Sundar, Chao-Han Huck Yang, David M. Chan, Shalini Ghosh, Venkatesh Ravichandran, Phani Sankar Nidadavolu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 5 pages, 1 figure, ICASSP 2024 Workshop on Self-supervision in Audio, Speech and Beyond

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09138 2024-02-07 cs.CV 79%

An Open-source Benchmark of Deep Learning Models for Audio-visual Apparent and Self-reported Personality Recognition

Rongfan Liao, Siyang Song, Hatice Gunes

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Affective Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.18084 2024-02-01 cs.CV cs.RO 79%

Binding Touch to Everything: Learning Unified Multimodal Tactile Representations

Fengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park, Daniel Wang, Yiming Dou, Ziyao Zeng, Xien Chen, Rit Gangopadhyay, Andrew Owens, Alex Wong

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10687 2024-02-01 eess.AS cs.SD 79%

MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis

Wenhao Guan, Yishuang Li, Tao Li, Hukai Huang, Feng Wang, Jiayan Lin, Lingyan Huang, Lin Li, Qingyang Hong

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Accepted at AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05669 2024-01-17 cs.CV 79%

Multi-Modal Gaze Following in Conversational Scenarios

Yuqi Hou, Zhongqun Zhang, Nora Horanyi, Jaewon Moon, Yihua Cheng, Hyung Jin Chang

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02746 2024-01-08 cs.CV 79%

Reading Between the Frames: Multi-Modal Depression Detection in Videos from Non-Verbal Cues

David Gimeno-Gómez, Ana-Maria Bucur, Adrian Cosma, Carlos-David Martínez-Hinarejos, Paolo Rosso

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at 46th European Conference on Information Retrieval (ECIR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00430 2024-01-04 cs.AI 79%

Brain-Conditional Multimodal Synthesis: A Survey and Taxonomy

Weijian Mai, Jian Zhang, Pengfei Fang, Zhijun Zhang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00424 2024-01-02 cs.CL 79%

SDIF-DA: A Shallow-to-Deep Interaction Framework with Data Augmentation for Multi-modal Intent Detection

Shijue Huang, Libo Qin, Bingbing Wang, Geng Tu, Ruifeng Xu

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17262 2024-01-01 cs.CL cs.LG 79%

Multimodal Classification of Teaching Activities from University Lecture Recordings

Oscar Sapena, Eva Onaindia

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 18 pages

Journal ref Appl. Sci. 2022, 12, 4785

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08288 2023-12-20 cs.CV 79%

Improving Audio-Visual Segmentation with Bidirectional Generation

Dawei Hao, Yuxin Mao, Bowen He, Xiaodong Han, Yuchao Dai, Yiran Zhong

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments AAAI Camera Ready. Dawei Hao and Yuxin Mao contribute equality to this paper. Yiran Zhong is the corresponding author. The code will be released at https://github.com/OpenNLPLab/AVS-bidirectional

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12558 2023-12-19 cs.CV 79%

Hyperbolic Audio-visual Zero-shot Learning

Jie Hong, Zeeshan Hayder, Junlin Han, Pengfei Fang, Mehrtash Harandi, Lars Petersson

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08850 2023-12-15 cs.SD eess.AS 79%

Hourglass-AVSR: Down-Up Sampling-based Computational Efficiency Model for Audio-Visual Speech Recognition

Fan Yu, Haoxu Wang, Ziyang Ma, Shiliang Zhang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted by ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.11871 2023-12-14 cs.CL 79%

Cem Mil Podcasts: A Spoken Portuguese Document Corpus For Multi-modal, Multi-lingual and Multi-Dialect Information Access Research

Ekaterina Garmash, Edgar Tanaka, Ann Clifton, Joana Correia, Sharmistha Jat, Winstead Zhu, Rosie Jones, Jussi Karlgren

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments 12 pages, 1 figure

Journal ref Volume 14163 of Lecture Notes in Computer Science, pages 48-59, Springer, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04131 2023-12-08 eess.AS cs.SD 79%

Joint Training or Not: An Exploration of Pre-trained Speech Models in Audio-Visual Speaker Diarization

Huan Zhao, Li Zhang, Yue Li, Yannan Wang, Hongji Wang, Wei Rao, Qing Wang, Lei Xie

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03632 2023-12-07 cs.SD cs.LG eess.AS 79%

Multimodal Data and Resource Efficient Device-Directed Speech Detection with Large Foundation Models

Dominik Wagner, Alexander Churchill, Siddharth Sigtia, Panayiotis Georgiou, Matt Mirsamadi, Aarshee Mishra, Erik Marchi

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01568 2023-12-05 cs.HC cs.SD eess.AS 79%

Multimodal Speech Emotion Recognition Using Modality-specific Self-Supervised Frameworks

Rutherford Agbeshi Patamia, Paulo E. Santos, Kingsley Nketia Acheampong, Favour Ekong, Kwabena Sarpong, She Kun

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17177 2023-11-30 cs.CV 79%

THInImg: Cross-modal Steganography for Presenting Talking Heads in Images

Lin Zhao, Hongxuan Li, Xuefei Ning, Xinru Jiang

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16471 2023-11-29 cs.CV 79%

A Unified Framework for Multimodal, Multi-Part Human Motion Synthesis

Zixiang Zhou, Yu Wan, Baoyuan Wang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments 19 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16446 2023-11-29 cs.CV 79%

Centre Stage: Centricity-based Audio-Visual Temporal Action Detection

Hanyuan Wang, Majid Mirmehdi, Dima Damen, Toby Perrett

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted to VUA workshop at BMVC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07860 2023-11-29 cs.SD cs.LG eess.AS 79%

EPG2S: Speech Generation and Speech Enhancement based on Electropalatography and Audio Signals using Multimodal Learning

Li-Chin Chen, Po-Hsun Chen, Richard Tzong-Han Tsai, Yu Tsao

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Accepted By IEEE Signal Processing Letter

Journal ref IEEE Signal Processing Letters, vol. 29, p. 2582-2586, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13165 2023-11-23 cs.AI 79%

Multimodal Large Language Models: A Survey

Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, Philip S. Yu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments IEEE BigData 2023. 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.00175 2023-11-22 eess.AS 79%

Multimodal Urban Sound Tagging with Spatiotemporal Context

Jisheng Bai, Jianfeng Chen, Mou Wang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11892 2023-11-21 cs.MM 79%

Multimodal Characterization of Emotion within Multimedia Space

Dayo Samuel Banjo, Connice Trimmingham, Niloofar Yousefi, Nitin Agarwal

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments 8 pages, Published in International Conference on Computers and Computation (COMPUTE 2022), November 03-04, 2022, San Francisco, United States

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10455 2023-11-21 eess.AS cs.SD 79%

Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement

Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Submmited to IEEE/ACM Transactions on Audio, Speech and Language Processing. arXiv admin note: text overlap with arXiv:2305.14933

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14933 2023-11-21 eess.AS cs.SD 79%

Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement through Knowledge Distillation

Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Published in InterSpeech 2023

Journal ref Proc. INTERSPEECH 2023, 844-848 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00641 2023-11-21 cs.CV 79%

RegBN: Batch Normalization of Multimodal Data with Regularization

Morteza Ghahremani, Christian Wachinger

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref Conference on Neural Information Processing Systems (NeurIPS 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06532 2023-11-14 cs.CL 79%

Added Toxicity Mitigation at Inference Time for Multimodal and Massively Multilingual Translation

Marta R. Costa-jussà, David Dale, Maha Elbayad, Bokai Yu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05190 2023-11-10 cs.CV 79%

Audio-visual Saliency for Omnidirectional Videos

Yuxin Zhu, Xilei Zhu, Huiyu Duan, Jie Li, Kaiwei Zhang, Yucheng Zhu, Li Chen, Xiongkuo Min, Guangtao Zhai

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments 13 pages, 5 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00595 2023-10-31 cs.CV 79%

Revisit Weakly-Supervised Audio-Visual Video Parsing from the Language Perspective

Yingying Fan, Yu Wu, Bo Du, Yutian Lin

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17568 2023-10-27 cs.HC cs.CL cs.RO 79%

Navigating to Success in Multi-Modal Human-Robot Collaboration: Analysis and Corpus Release

Stephanie M. Lukin, Kimberly A. Pollard, Claire Bonial, Taylor Hudson, Ron Arstein, Clare Voss, David Traum

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments 7 pages, 3 figures

Journal ref Proceedings of the 2023 IEEE Robot and Human Interactive Communication Conference

详情

展开后加载摘要…

URL PDF HTML 收藏