arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4559 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4559 篇

2503.22197 2025-03-31 cs.CV 79%

Extremely Simple Out-of-distribution Detection for Audio-visual Generalized Zero-shot Learning

Yang Liu, Xun Zhang, Jiale Du, Xinbo Gao, Jungong Han

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17453 2025-03-25 cs.CV 79%

Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition

Ran Liu, Fengyu Zhang, Cong Yu, Longjiang Yang, Zhuofan Wen, Siyuan Zhang, Hailiang Yao, Shun Chen, Zheng Lian, Bin Liu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12605 2025-03-25 cs.CV 79%

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Yaoting Wang, Shengqiong Wu, Yuecheng Zhang, Shuicheng Yan, Ziwei Liu, Jiebo Luo, Hao Fei

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments Survey, working under progress; 12 figures, 4 tables, 44 pages; Resource at https://github.com/yaotingwangofficial/Awesome-MCoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16454 2025-03-24 cs.HC cs.AI 79%

An Audio-Visual Fusion Emotion Generation Model Based on Neuroanatomical Alignment

Haidong Wang, Qia Shan, JianHua Zhang, PengFei Xiao, Ao Liu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13068 2025-03-18 cs.CV 79%

Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

Henghui Du, Guangyao Li, Chang Zhou, Chunjie Zhang, Alan Zhao, Di Hu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07862 2025-03-18 cs.CL 79%

cantnlp@DravidianLangTech2025: A Bag-of-Sounds Approach to Multimodal Hate Speech Detection

Sidney Wong, Andrew Li

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted Fifth Workshop on Speech and Language Technologies for Dravidian Languages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09511 2025-03-13 cs.CL 79%

TRACE: Real-Time Multimodal Common Ground Tracking in Situated Collaborative Dialogues

Hannah VanderHoeven, Brady Bhalla, Ibrahim Khebour, Austin Youngren, Videep Venkatesha, Mariah Bradford, Jack Fitzgerald, Carlos Mabrey, Jingxuan Tu, Yifan Zhu, Kenneth Lai, Changsoo Jung, James Pustejovsky, Nikhil Krishnaswamy

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 11 pages, 4 tables, 4 figures, to appear at NAACL 2025 Demos program, Albuquerque, NM, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08335 2025-03-12 cs.CV 79%

Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos

Soumya Shamarao Jahagirdar, Jayasree Saha, C V Jawahar

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments CVIP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05379 2025-03-11 cs.LG cs.CV 79%

R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Jiaxing Zhao, Xihan Wei, Liefeng Bo

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03190 2025-03-11 cs.LG cs.HC eess.AS eess.IV 79%

Multimodal Machine Learning Can Predict Videoconference Fluidity and Enjoyment

Andrew Chang, Viswadruth Akkaraju, Ray McFadden Cogliano, David Poeppel, Dustin Freeman

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05008 2025-03-10 cs.MM 79%

Enhancing Video Music Recommendation with Transformer-Driven Audio-Visual Embeddings

Shimiao Liu, Alexander Lerch

专题命中 音频语音多模态 :audio-visual(title);cross-modal(abstract);分类 cs.MM

Comments 2024 IEEE 5th International Symposium on the Internet of Sounds (IS2), Erlangen, Germany, 2024, pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18501 2025-03-04 eess.AS cs.SD 79%

Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention

Ruijie Tao, Xinyuan Qian, Yidi Jiang, Junjie Li, Jiadong Wang, Haizhou Li

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15711 2025-02-25 cs.IR cs.MM 79%

A Survey on Multimodal Recommender Systems: Recent Advances and Future Directions

Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li, Wei Wang, Xiping Hu, Steven Hoi, Edith Ngai

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13133 2025-02-19 cs.CV 79%

AV-Flow: Transforming Text to Audio-Visual Human-like Interactions

Aggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhoefer, Dimitris Samaras, Alexander Richard

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12771 2025-02-19 cs.CL q-bio.NC 79%

Mind the Gap: Aligning the Brain with Language Models Requires a Nonlinear and Multimodal Approach

Danny Dongyeop Han, Yunju Cho, Jiook Cha, Jay-Yoon Lee

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09928 2025-02-19 cs.SD eess.AS 79%

Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's Disease Identification

Yifan Gao, Long Guo, Hong Liu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08862 2025-02-14 eess.AS 79%

Predicting Cognitive Decline: A Multimodal AI Approach to Dementia Screening from Speech

Lei Chi, Arav Sharma, Ari Gebhardt, Joseph T. Colonel

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Submitted to IEEE ICAD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00832 2025-02-10 cs.AI cs.CY 79%

Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education

Arne Bewersdorff, Christian Hartmann, Marie Hornberger, Kathrin Seßler, Maria Bannert, Enkelejda Kasneci, Gjergji Kasneci, Xiaoming Zhai, Claudia Nerdel

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments revised version 2. September 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13416 2025-02-04 cs.LG cs.AI cs.RO 79%

M3PT: A Transformer for Multimodal, Multi-Party Social Signal Prediction with Person-aware Blockwise Attention

Yiming Tang, Abrar Anwar, Jesse Thomason

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12674 2025-01-23 eess.AS cs.SD 79%

EmoTech: A Multi-modal Speech Emotion Recognition Using Multi-source Low-level Information with Hybrid Recurrent Network

Shamin Bin Habib Avro, Taieba Taher, Nursadul Mamun

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11468 2025-01-22 eess.AS cs.SD 79%

LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations

Soumya Dutta, Sriram Ganapathy

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments ICASSP 2025; 5 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09506 2025-01-22 cs.LG cs.SD eess.AS eess.IV 79%

Multimodal Marvels of Deep Learning in Medical Diagnosis: A Comprehensive Review of COVID-19 Detection

Md Shofiqul Islam, Khondokar Fida Hasan, Hasibul Hossain Shajeeb, Humayan Kabir Rana, Md Saifur Rahmand, Md Munirul Hasan, AKM Azad, Ibrahim Abdullah, Mohammad Ali Moni

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 43 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08124 2025-01-15 eess.AS cs.SD 79%

Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech

Mareike Daeglau, Juergen Otten, Giso Grimm, Bojana Mirkovic, Volker Hohmann, Stefan Debener

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06530 2025-01-14 eess.AS cs.SD 79%

Multi-modal Speech Enhancement with Limited Electromyography Channels

Fuyuan Feng, Longting Xu, Rohan Kumar Das

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00481 2025-01-09 eess.AS cs.SD 79%

DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module

Xinyu Wang, Haotian Jiang, Haolin Huang, Yu Fang, Mengjie Xu, Qian Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted to ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04138 2025-01-09 cs.CL 79%

"Yeah Right!" -- Do LLMs Exhibit Multimodal Feature Transfer?

Benjamin Reichman, Kartik Talamadupula

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01401 2025-01-03 eess.AS 79%

VoiceVector: Multimodal Enrolment Vectors for Speaker Separation

Akam Rahimi, Triantafyllos Afouras, Andrew Zisserman

专题命中 音频语音多模态 :multimodal(title);audio-visual(abstract);分类 eess.AS

Journal ref 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00029 2025-01-03 cs.CL cs.IR cs.LG 79%

A Breadth-First Catalog of Text Processing, Speech Processing and Multimodal Research in South Asian Languages

Pranav Gupta

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20872 2025-01-03 cs.CV 79%

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing

Langyu Wang, Bingke Zhu, Yingying Chen, Jinqiao Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20467 2024-12-31 cs.CL 79%

Utilizing Multimodal Data for Edge Case Robust Call-sign Recognition and Understanding

Alexander Blatt, Dietrich Klakow

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏