arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4557 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4557 篇

2508.00760 2025-08-04 cs.CL cs.AI 81%

MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations

Qiyao Xue, Yuchen Dou, Ryan Shi, Xiang Lorraine Li, Wei Gao

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22934 2025-08-01 cs.CL cs.AI 81%

Deep Learning Approaches for Multimodal Intent Recognition: A Survey

Jingwei Zhao, Yuhua Wen, Qifei Li, Minchi Hu, Yingying Zhou, Jingyao Xue, Junyang Wu, Yingming Gao, Zhengqi Wen, Jianhua Tao, Ya Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Submitted to ACM Computing Surveys

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19956 2025-07-29 cs.CV cs.AI q-bio.NC 81%

Predicting Brain Responses To Natural Movies With Multimodal LLMs

Cesar Kadir Torrico Villanueva, Jiaxin Cindy Tu, Mihir Tripathy, Connor Lane, Rishab Iyer, Paul S. Scotti

机构 * Medical AI Research Center (MedARC)(医学人工智能研究中心(MedARC)) Psychological and Brain Sciences, Dartmouth College(心理学与脑科学系,达特茅斯学院) Core for Advanced Magnetic Resonance Imaging (CAMRI), Baylor College of Medicine(先进磁共振成像核心(CAMRI),贝勒医学院) Sophont Princeton Neuroscience Institute(普林斯顿神经科学研究所)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Code available at https://github.com/MedARC-AI/algonauts2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08714 2025-07-29 cs.CV cs.AI 81%

Versatile Multimodal Controls for Expressive Talking Human Animation

Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li, Zixin Zhu, Sanping Zhou, Ming Yang, Le Wang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University, Ant Group(人机混合增强智能国家级实验室,西安交通大学,蚂蚁集团) National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University(人机混合增强智能国家级实验室,西安交通大学) University at Buffalo(布法罗大学)

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM MM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23822 2025-07-24 cs.CL cs.MM 81%

Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction

Mai Ali, Christopher Lucasius, Tanmay P. Patel, Madison Aitken, Jacob Vorstman, Peter Szatmari, Marco Battaglia, Deepa Kundur

机构 * The Edward S. Rogers Sr. Department of Electrical and Computer Engineering, University of Toronto, Toronto, Canada(电气与计算机工程系,多伦多大学) Division of Engineering Science, University of Toronto, Toronto, Canada(工程科学系,多伦多大学) Cundill Centre for Child and Youth Depression, Centre for Addiction and Mental Health, Toronto, Canada(儿童与青少年抑郁研究中心,成瘾与心理健康中心) Department of Psychology, York University, Toronto, Canada(心理学系,约克大学) The Hospital for Sick Children, Toronto, ON, Canada(多伦多儿童医院) Department of Psychiatry, University of Toronto, Toronto, Canada(精神病学系,多伦多大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments 6 pages, 1 figure, 3 tables. The corresponding author is Mai Ali (maia dot ali at mail dot utoronto dot ca). Christopher Lucasius and Tanmay P. Patel contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14579 2025-07-22 cs.CL cs.AI 81%

Exploring Human-AI Complementarity in CPS Diagnosis Using Unimodal and Multimodal BERT Models

Kester Wong, Sahan Bulathwela, Mutlu Cukurova

机构 * UCL Knowledge Lab, Institute of Education, University College London, UK(伦敦大学学院教育研究所知识实验室) UCL Centre for Artificial Intelligence, Department of Computer Science, University College London, UK(伦敦大学学院人工智能中心计算机科学系)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to appear in the workshop proceedings for the HEXED'25 workshop in the 26th International Conference on Artificial Intelligence in Education 2025 (AIED 2025), 22 July 2025, Palermo, Italy. 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07046 2025-07-21 cs.AI cs.CV 81%

CorMulT: A Semi-supervised Modality Correlation-aware Multimodal Transformer for Sentiment Analysis

Yangmin Li, Ruiqi Zhu, Wengen Li

机构 * Information Networking Institute at Carnegie Mellon University (CMU)(卡内基梅隆大学信息网络研究所) School of Computer Science and Technology at Tongji University(同济大学计算机科学与技术学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13364 2025-07-21 cs.CV cs.AI 81%

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

Siddharth Srivastava, Gaurav Sharma

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12723 2025-07-18 cs.SD cs.MM eess.AS 81%

Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries

Minyoung Kim, Sehwan Park, Sungmin Cha, Paul Hongsuck Seo

机构 * Korea University(韩国大学) New York University(纽约大学)

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.MM、eess.AS

Comments 5 pages, 2 figures, Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11967 2025-07-17 cs.CV eess.AS eess.IV 81%

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

Yuchi Ishikawa, Shota Nakada, Hokuto Munakata, Kazuhiro Saito, Tatsuya Komatsu, Yoshimitsu Aoki

机构 * LY Corporation(LY公司) Keio University(庆应大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09862 2025-07-15 cs.CV eess.AS 81%

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

Youliang Zhang, Zhaoyang Li, Duomin Wang, Jiahe Zhang, Deyu Zhou, Zixin Yin, Xili Dai, Gang Yu, Xiu Li

机构 * Tsinghua University(清华大学) StepFun The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07270 2025-07-11 cs.SD cs.MM eess.AS 81%

Audio-Visual Speech Separation via Bottleneck Iterative Network

Sidong Zhang, Shiv Shankar, Trang Nguyen, Andrea Fanelli, Madalina Fiterau

机构 * Manning College of Information \& Computer Sciences, University of Massachusetts Amherst, Amherst, U.S. Dolby Laboratories, San Francisco, U.S.

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.MM、eess.AS

Comments Accepted to the 42nd International Conference on Machine Learning Workshop on Machine Learning for Audio

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23714 2025-07-01 cs.CV cs.CL 81%

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization

Md Moinul Islam, Sofoklis Kakouros, Janne Heikkilä, Mourad Oussalah

机构 * Center for Machine Vision and Signal Analysis, Faculty of ITEE, University of Oulu, Finland(机器视觉与信号分析中心,ITEE学院,奥卢大学,芬兰) Center for Machine Vision(机器视觉与信号分析中心) Signal Analysis, Faculty of ITEE, University of Oulu, Finland(信号分析,ITEE学院,奥卢大学,芬兰)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to HHAI WS 2025: Workshops at the Fourth International Conference on Hybrid Human-Artificial Intelligence (HHAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05460 2025-07-01 cs.DC cs.AI cs.CV cs.LG 81%

Efficiently Serving Large Multimodal Models Using EPD Disaggregation

Gursimran Singh, Xinglu Wang, Yifan Hu, Timothy Yu, Linzi Xing, Wei Jiang, Zhefeng Wang, Xiaolong Bai, Yi Li, Ying Xiong, Yong Zhang, Zhenan Fan

机构 * Huawei Technologies Canada, BC, Canada(华为加拿大技术有限公司) Simon Fraser University, BC, Canada(西蒙弗雷泽大学) Huawei Cloud, China(华为云)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 17 pages, 12 figures, 9 tables

Journal ref International Conference on Machine Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17694 2025-06-25 cs.CV cs.SD eess.AS 81%

SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification

Gnana Praveen Rajasekhar, Jahangir Alam

机构 * Computer Research Institute of Montreal (CRIM), Canada(蒙特利尔计算机研究院(CRIM))

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09792 2025-06-17 cs.SD cs.LG cs.MM eess.AS 81%

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

Wenxuan Wu, Shuai Wang, Xixin Wu, Helen Meng, Haizhou Li

机构 * Department of SEEM(SEEM系) SRIBD, School of Data Science(数据科学学院) School of Intelligence Science and Technology(智能科学与技术学院) Department of ECE(电子工程系)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.MM、eess.AS

Comments Accepted by Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10010 2025-06-13 cs.MM cs.LG cs.SD eess.AS 81%

Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction

Von Ralph Dane Marquez Herbuela, Yukie Nagai

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08504 2025-06-11 cs.CL cs.AI cs.LG 81%

CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations

Divyaksh Shukla, Ritesh Baviskar, Dwijesh Gohil, Aniket Tiwari, Atul Shree, Ashutosh Modi

机构 * Indian Institute of Technology Kanpur (IIT Kanpur)(印度理工学院坎浦尔学院)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at ACL Findings 2025 (16 pages: 5 pages main content + 3 pages references + 8 pages appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03846 2025-06-03 cs.CV cs.AI 81%

GAME: Learning Multimodal Interactions via Graph Structures for Personality Trait Estimation

Kangsheng Wang, Yuhang Li, Chengwei Ye, Yufei Lin, Huanzhen Zhang, Bohan Hu, Linuo Xu, Shuyan Liu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments The article contains serious scientific errors and cannot be corrected by updating the preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24493 2025-06-02 cs.AI cs.SD eess.AS 81%

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge

Xin Jing, Jiadong Wang, Iosif Tsangko, Andreas Triantafyllopoulos, Björn W. Schuller

机构 * Technical University of Munich(慕尼黑技术大学) Imperial College London(伦敦帝国理工学院) CHI – Chair of Health Informatics(健康信息学系) Munich Centre for Machine Learning(慕尼黑机器学习中心) GLAM – Group on Language, Audio, & Music(语言、音频与音乐小组)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20635 2025-05-28 eess.AS cs.AI cs.SD 81%

Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction

Zexu Pan, Shengkui Zhao, Tingting Wang, Kun Zhou, Yukun Ma, Chong Zhang, Bin Ma

机构 * Alibaba Group(阿里巴巴集团) Nanjing University of Posts and Telecommunications(南京邮电大学) Tongyi Lab(通义实验室)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI、eess.AS

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18453 2025-05-27 cs.SD cs.AI eess.AS 81%

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt

Zhichao Wu, Yueteng Kang, Songjun Cao, Long Ma, Qiulin Li, Qun Yang

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) TencentChina(腾讯中国) Youtu Lab(你图实验室)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI、eess.AS

Comments Accepted by InterSpeech

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16735 2025-05-26 eess.AS cs.AI 81%

Adversarial Deep Metric Learning for Cross-Modal Audio-Text Alignment in Open-Vocabulary Keyword Spotting

Youngmoon Jung, Yong-Hyeok Lee, Myunghun Jung, Jaeyoung Roh, Chang Woo Han, Hoon-Young Cho

机构 * Samsung Research(三星研究院)

专题命中 音频语音多模态 :cross-modal(title);multi-modal(abstract);分类 cs.AI、eess.AS

Comments 5 pages, 1 figure, Accepted at Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10447 2025-05-22 eess.AS cs.CL cs.LG 81%

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

Sungnyun Kim, Kangwook Jang, Sangmin Bae, Sungwoo Cho, Se-Young Yun

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CL、eess.AS

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18799 2025-04-29 cs.MM cs.SD eess.AS 81%

A Survey on Multimodal Music Emotion Recognition

Rashini Liyanarachchi, Aditya Joshi, Erik Meijering

机构 * University of New South Wales(新南威尔士大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18547 2025-04-22 cs.CR cs.AI cs.MA cs.MM 81%

Steganography Beyond Space-Time with Chain of Multimodal AI

Ching-Chun Chang, Isao Echizen

机构 * Information and Society Research Division, National Institute of Informatics(信息与社会研究部,日本信息机构) Graduate School of Information Science and Technology, University of Tokyo(东京大学信息科学与技术研究生院) School of Multidisciplinary Sciences, Graduate University for Advanced Studies (SOKENDAI)(多学科科学学院,研究生高级研究大学(SOKENDAI))

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

Journal ref Scientific Reports, vol. 15, no. 1, Article 12908, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13535 2025-04-21 cs.SD cs.MM eess.AS 81%

MusFlow: Multimodal Music Generation via Conditional Flow Matching

Jiahao Song, Yuzhao Wang

机构 * School of Mathematics and Statistics, Shanxi University(山西大学数学与统计学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14778 2025-04-15 cs.MM cs.SD eess.AS 81%

Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions

Jinzheng Zhao, Yong Xu, Xinyuan Qian, Davide Berghi, Peipei Wu, Meng Cui, Jianyuan Sun, Philip J. B. Jackson, Wenwu Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00221 2025-04-09 cs.HC cs.AI cs.CV 81%

GazeLLM: Multimodal LLMs incorporating Human Visual Attention

Jun Rekimoto

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref Augmented Humans 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15029 2025-03-25 cs.CL cs.AI 81%

Enhancing Multimodal Sentiment Analysis for Missing Modality through Self-Distillation and Unified Modality Cross-Attention

Yuzhe Weng, Haotian Wang, Tian Gao, Kewei Li, Shutong Niu, Jun Du

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏