arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4559 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4559 篇

2308.11596 2023-10-26 cs.CL 79%

SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Seamless Communication, Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, Christopher Klaiber, Pengwei Li, Daniel Licht, Jean Maillard, Alice Rakotoarison, Kaushik Ram Sadagopan, Guillaume Wenzek, Ethan Ye, Bapi Akula, Peng-Jen Chen, Naji El Hachem, Brian Ellis, Gabriel Mejia Gonzalez, Justin Haaheim, Prangthip Hansanti, Russ Howes, Bernie Huang, Min-Jae Hwang, Hirofumi Inaguma, Somya Jain, Elahe Kalbassi, Amanda Kallet, Ilia Kulikov, Janice Lam, Daniel Li, Xutai Ma, Ruslan Mavlyutov, Benjamin Peloquin, Mohamed Ramadan, Abinesh Ramakrishnan, Anna Sun, Kevin Tran, Tuan Tran, Igor Tufanov, Vish Vogeti, Carleigh Wood, Yilin Yang, Bokai Yu, Pierre Andrews, Can Balioglu, Marta R. Costa-jussà, Onur Celebi, Maha Elbayad, Cynthia Gao, Francisco Guzmán, Justine Kao, Ann Lee, Alexandre Mourachko, Juan Pino, Sravya Popuri, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, Paden Tomasello, Changhan Wang, Jeff Wang, Skyler Wang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11073 2023-10-17 cs.CV 79%

Audio-Visual Class-Incremental Learning

Weiguo Pian, Shentong Mo, Yunhui Guo, Yapeng Tian

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07478 2023-10-13 cs.AI 79%

Multimodal Graph Learning for Generative Tasks

Minji Yoon, Jing Yu Koh, Bryan Hooi, Ruslan Salakhutdinov

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11059 2023-10-10 eess.AS cs.SD 79%

Deep Complex U-Net with Conformer for Audio-Visual Speech Enhancement

Shafique Ahmed, Chia-Wei Chen, Wenze Ren, Chin-Jou Li, Ernie Chu, Jun-Cheng Chen, Amir Hussain, Hsin-Min Wang, Yu Tsao, Jen-Cheng Hou

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03724 2023-10-09 cs.CL 79%

Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer

Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CL

Journal ref Proceedings of Interspeech 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11081 2023-09-21 cs.CV 79%

Dense 2D-3D Indoor Prediction with Sound via Aligned Cross-Modal Distillation

Heeseung Yun, Joonil Na, Gunhee Kim

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Published to ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09709 2023-09-21 cs.CV 79%

CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video Segmentation

Kexin Li, Zongxin Yang, Lei Chen, Yi Yang, Jun Xiao

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08408 2023-09-18 cs.SD eess.AS 79%

Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-talker Speech

Junjie Li, Ruijie Tao, Zexu Pan, Meng Ge, Shuai Wang, Haizhou Li

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Submitted to ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05091 2023-09-12 cs.HC cs.MM 79%

SpeechMirror: A Multimodal Visual Analytics System for Personalized Reflection of Online Public Speaking Effectiveness

Zeyuan Huang, Qiang He, Kevin Maher, Xiaoming Deng, Yu-Kun Lai, Cuixia Ma, Sheng-feng Qin, Yong-Jin Liu, Hongan Wang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments Main paper (11 pages, 6 figures) and Supplemental document (11 pages, 11 figures). Accepted by VIS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14274 2023-08-29 cs.MM 79%

Parameter-Efficient Transfer Learning for Audio-Visual-Language Tasks

Hongye Liu, Xianhai Xie, Yang Gao, Size Li, Zhou YU

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14052 2023-08-29 cs.CV 79%

MM-AU:Towards Multimodal Understanding of Advertisement Videos

Digbalay Bose, Rajat Hebbar, Tiantian Feng, Krishna Somandepalli, Anfeng Xu, Shrikanth Narayanan

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11466 2023-08-24 cs.CL 79%

SONAR: Sentence-Level Multimodal and Language-Agnostic Representations

Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10306 2023-08-22 cs.CV 79%

Omnidirectional Information Gathering for Knowledge Transfer-based Audio-Visual Navigation

Jinyu Chen, Wenguan Wang, Si Liu, Hongsheng Li, Yi Yang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12300 2023-08-21 cs.SD eess.AS 79%

A Multimodal Prototypical Approach for Unsupervised Sound Classification

Saksham Singh Kushwaha, Magdalena Fuentes

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Accepted to INTERSPEECH 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04778 2023-08-10 cs.AI 79%

Multi-modal Multi-view Clustering based on Non-negative Matrix Factorization

Yasser Khalafaoui, Nistor Grozavu, Basarab Matei, Laurent-Walter Goix

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

Journal ref 2022 IEEE Symposium Series on Computational Intelligence (SSCI), Dec 2022, Singapore, Singapore. pp.1386-1391

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04767 2023-08-10 cs.CV cs.AI cs.MM cs.SD eess.AS 79%

Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization

Tianyu Liu, Peng Zhang, Wei Huang, Yufei Zha, Tao You, Yanning Zhang

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.04930 2023-08-10 cs.LG cs.CV cs.RO 79%

Multimodal Multi-User Surface Recognition with the Kernel Two-Sample Test

Behnam Khojasteh, Friedrich Solowjow, Sebastian Trimpe, Katherine J. Kuchenbecker

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01864 2023-08-02 cs.SD cs.LG eess.AS 79%

Unsupervised Improvement of Audio-Text Cross-Modal Representations

Zhepei Wang, Cem Subakan, Krishna Subramani, Junkai Wu, Tiago Tavares, Fabio Ayres, Paris Smaragdis

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 eess.AS

Comments Accepted to WASPAA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15400 2023-07-31 cs.SD eess.AS 79%

The FlySpeech Audio-Visual Speaker Diarization System for MISP Challenge 2022

Li Zhang, Huan Zhao, Yue Li, Bowen Pang, Yannan Wang, Hongji Wang, Wei Rao, Qing Wang, Lei Xie

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00610 2023-07-28 cs.LG cs.CL cs.SI 79%

Fraunhofer SIT at CheckThat! 2023: Mixing Single-Modal Classifiers to Estimate the Check-Worthiness of Multi-Modal Tweets

Raphael Frick, Inna Vogel

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments 8 pages

Journal ref CLEF 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13829 2023-07-27 cs.CL cs.SI 79%

ARC-NLP at Multimodal Hate Speech Event Detection 2023: Multimodal Methods Boosted by Ensemble Learning, Syntactical and Entity Features

Umitcan Sahin, Izzet Emre Kucukkaya, Oguzhan Ozcelik, Cagri Toraman

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Submitted to CASE at RANLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12236 2023-07-25 cs.LG cs.CV 79%

Multi-Modal Machine Learning for Assessing Gaming Skills in Online Streaming: A Case Study with CS:GO

Longxiang Zhang, Wenping Wang

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02609 2023-07-07 cs.CV 79%

MRecGen: Multimodal Appropriate Reaction Generator

Jiaqi Xu, Cheng Luo, Weicheng Xie, Linlin Shen, Xiaofeng Liu, Lu Liu, Hatice Gunes, Siyang Song

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09947 2023-06-19 eess.AS 79%

Knowledge Distillation for Efficient Audio-Visual Video Captioning

Özkan Çaylı, Xubo Liu, Volkan Kılıç, Wenwu Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments European Signal Processing Conference (EUSIPCO 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12831 2023-06-13 eess.AS cs.SD 79%

Target Active Speaker Detection with Audio-visual Cues

Yidi Jiang, Ruijie Tao, Zexu Pan, Haizhou Li

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted to INTERSPEECH2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06652 2023-06-13 cs.SD eess.AS 79%

Audio-Visual Mandarin Electrolaryngeal Speech Voice Conversion

Yung-Lun Chien, Hsin-Hao Chen, Ming-Chi Yen, Shu-Wei Tsai, Hsin-Min Wang, Yu Tsao, Tai-Shih Chi

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 eess.AS

Comments Accepted to INTERSPEECH 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06495 2023-06-13 eess.AS cs.SD 79%

Audio-Visual Speech Enhancement With Selective Off-Screen Speech Extraction

Tomoya Yoshinaga, Keitaro Tanaka, Shigeo Morishima

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted by EUSIPCO 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06036 2023-06-12 cs.AI 79%

SNeL: A Structured Neuro-Symbolic Language for Entity-Based Multimodal Scene Understanding

Silvan Ferreira, Allan Martins, Ivanovitch Silva

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02625 2023-06-06 cs.SD eess.AS 79%

Rethinking the visual cues in audio-visual speaker extraction

Junjie Li, Meng Ge, Zexu pan, Rui Cao, Longbiao Wang, Jianwu Dang, Shiliang Zhang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted in Interspeech 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01432 2023-06-05 eess.AS cs.LG 79%

Audio-Visual Speech Enhancement with Score-Based Generative Models

Julius Richter, Simone Frintrop, Timo Gerkmann

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Submitted to ITG Conference on Speech Communication

详情

展开后加载摘要…

URL PDF HTML 收藏