arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4559 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4559 篇

2412.19563 2024-12-30 cs.CV 79%

Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing

Yongbiao Gao, Xiangcheng Sun, Guohua Lv, Deng Yu, Sijiu Niu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16771 2024-12-24 cs.CV 79%

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization

Tan-Hanh Pham, Hoang-Nam Le, Phu-Vinh Nguyen, Chris Ngo, Truong-Son Hy

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13387 2024-12-19 eess.AS cs.SD 79%

Deep Speech Synthesis from Multimodal Articulatory Representations

Peter Wu, Bohan Yu, Kevin Scheck, Alan W Black, Aditi S. Krishnapriyan, Irene Y. Chen, Tanja Schultz, Shinji Watanabe, Gopala K. Anumanchipalli

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10103 2024-12-16 cs.CL 79%

AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation

Xiyuan Gao, Shubhi Bansal, Kushaan Gowda, Zhu Li, Shekhar Nayak, Nagendra Kumar, Matt Coler

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments This is a preprint version of the paper, submitted and under review at the IEEE Transactions on Affective Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09317 2024-12-13 cs.SD cs.AI cs.CV cs.MM eess.AS 79%

Multimodal Sentiment Analysis based on Video and Audio Inputs

Antonio Fernandez, Suzan Awinat

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.AI、cs.MM

Comments Presented as a full paper in the 15th International Conference on Emerging Ubiquitous Systems and Pervasive Networks (EUSPN 2024) October 28-30, 2024, Leuven, Belgium

Journal ref Procedia Computer Science, Volume 251, 2024, Pages 41-48, ISSN 1877-0509

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08529 2024-12-12 cs.CL 79%

TECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction

Quynh-Mai Thi Nguyen, Lan-Nhi Thi Nguyen, Cam-Van Thi Nguyen

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at PACLIC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16213 2024-12-12 cs.CV 79%

SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context

Jungang Li, Sicheng Tao, Yibo Yan, Xiaojie Gu, Haodong Xu, Xu Zheng, Yuanhuiyi Lyu, Linfeng Zhang, Xuming Hu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments The publication has some processing errors (language short-cuts in synthetic data are not avoided) that invalidate some of the conclusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05315 2024-12-10 cs.CL cs.CY 79%

Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor

Ashwin Baluja

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10234 2024-11-18 cs.HC cs.AI 79%

Generative AI in Multimodal User Interfaces: Trends, Challenges, and Cross-Platform Adaptability

J. Bieniek, M. Rahouti, D. C. Verma

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10880 2024-11-15 cs.CL 79%

Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR

Minghan Wang, Yuxia Wang, Thuy-Trang Vu, Ehsan Shareghi, Gholamreza Haffari

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06754 2024-11-12 cs.LG cs.AI 79%

Scaling Law Hypothesis for Multimodal Model

Qingyun Sun, Zhen Guo, PIN AI Team

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05603 2024-11-11 cs.CV 79%

Efficient Audio-Visual Fusion for Video Classification

Mahrukh Awan, Asmar Nadeem, Armin Mustafa

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments CVMP Short Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22112 2024-10-30 cs.MM 79%

Multimodal Semantic Communication for Generative Audio-Driven Video Conferencing

Haonan Tong, Haopeng Li, Hongyang Du, Zhaohui Yang, Changchuan Yin, Dusit Niyato

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments accepted by IEEE Wireless Communications Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21170 2024-10-29 cs.CV 79%

Joint Audio-Visual Idling Vehicle Detection with Streamlined Input Dependencies

Xiwen Li, Rehman Mohammed, Tristalee Mangin, Surojit Saha, Ross T Whitaker, Kerry E. Kelly, Tolga Tasdizen

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20916 2024-10-29 cs.CL 79%

NeuGPT: Unified multi-modal Neural GPT

Yiqian Yang, Yiqun Duan, Hyejeong Jo, Qiang Zhang, Renjing Xu, Oiwi Parker Jones, Xuming Hu, Chin-teng Lin, Hui Xiong

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20116 2024-10-29 cs.HC cs.AI 79%

Estuary: A Framework For Building Multimodal Low-Latency Real-Time Socially Interactive Agents

Spencer Lin, Basem Rizk, Miru Jun, Andy Artze, Caitlin Sullivan, Sharon Mozgai, Scott Fisher

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments To be published in ACM Intelligent Virtual Agents (IVA) 2024 [DOI: 10.1145/3652988.3696198] [ACM ISBN: 979-8-4007-0625-7/24/09]

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18882 2024-10-25 cs.CL 79%

A Survey of Multimodal Sarcasm Detection

Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanojia, Yu Kong, Marcos Zampieri

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Published in the Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence Survey Track. Pages 8020-8028

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03650 2024-10-22 cs.MM 79%

Towards Multimodal Emotional Support Conversation Systems

Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14917 2024-10-21 cs.CL 79%

With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models

Tyler Loakman, Yucheng Li, Chenghua Lin

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to EMNLP 2024 (Camera Ready)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13803 2024-10-18 cs.AI 79%

A Pattern to Align Them All: Integrating Different Modalities to Define Multi-Modal Entities

Gianluca Apriceno, Valentina Tamma, Tania Bailoni, Jacopo de Berardinis, Mauro Dragoni

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07757 2024-10-11 cs.CV 79%

MMHead: Towards Fine-grained Multi-modal 3D Facial Animation

Sijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan, Ziwei Liu, Guangtao Zhai

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ACMMM 2024. Project page: https://wsj-sjtu.github.io/MMHead/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06543 2024-10-10 cs.CR cs.SD eess.AS 79%

Gumbel Rao Monte Carlo based Bi-Modal Neural Architecture Search for Audio-Visual Deepfake Detection

Aravinda Reddy PN, Raghavendra Ramachandra, Krothapalli Sreenivasa Rao, Pabitra Mitra Vinod Rathod

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05608 2024-10-10 cs.CL 79%

Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond

Soyeon Caren Han, Feiqi Cao, Josiah Poon, Roberto Navigli

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at ACM-MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04819 2024-10-08 cs.CL 79%

MINER: Mining the Underlying Pattern of Modality-Specific Neurons in Multimodal Large Language Models

Kaichen Huang, Jiahao Huo, Yibo Yan, Kun Wang, Yutao Yue, Xuming Hu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19422 2024-10-02 cs.LG cs.AI stat.ML 79%

Identifiable Shared Component Analysis of Unpaired Multimodal Mixtures

Subash Timilsina, Sagar Shrestha, Xiao Fu

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14554 2024-10-01 eess.AS cs.SD 79%

Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing

Wenze Ren, Kuo-Hsuan Hung, Rong Chao, YouJin Li, Hsin-Min Wang, Yu Tsao

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments The 27th International Conference of the Oriental COCOSDA

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15741 2024-09-25 eess.AS cs.SD 79%

StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis

Zhiyong Chen, Xinnuo Li, Zhiqi Ai, Shugong Xu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments The 7th Chinese Conference on Pattern Recognition and Computer Vision PRCV 2024

Journal ref The 7th Chinese Conference on Pattern Recognition and Computer Vision PRCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13483 2024-09-23 cs.CL cs.IR 79%

A Multimodal Dense Retrieval Approach for Speech-Based Open-Domain Question Answering

Georgios Sidiropoulos, Evangelos Kanoulas

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05136 2024-09-18 cs.CL 79%

MHS-STMA: Multimodal Hate Speech Detection via Scalable Transformer-Based Multilevel Attention Framework

Anusha Chhabra, Dinesh Kumar Vishwakarma

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07256 2024-09-12 cs.CV 79%

MRAC Track 1: 2nd Workshop on Multimodal, Generative and Responsible Affective Computing

Shreya Ghosh, Zhixi Cai, Abhinav Dhall, Dimitrios Kollias, Roland Goecke, Tom Gedeon

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments ACM MM Workshop 2024. Workshop webpage: https://react-ws.github.io/2024/

详情

展开后加载摘要…

URL PDF HTML 收藏