arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4557 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4557 篇

2601.13098 2026-01-21 cs.HC 82%

Exploring the Impacts of Background Noise on Auditory Stimuli of Audio-Visual eHMIs for Hearing, Deaf, and Hard-of-Hearing People

探索背景噪声对听觉刺激在音频视觉eHMIs中的影响:面向听障、聋人及听力障碍人群

Wenge Xu, Foroogh Hajiseyedjavadi, Debargha Dey, Tram Thi Minh Tran, Mark Colley

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract)

AI总结 研究探讨了背景噪声对聋人及听力障碍者在音频视觉eHMI中听觉刺激的影响,发现嘈杂环境会损害行人穿越体验,而增加铃声或语音提示可改善体验。

Comments This is the author's version of the paper accepted at CHI Conference on Human Factors in Computing Systems (CHI '26), April 13-17, 2026, Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05879 2026-01-19 cs.HC 82%

Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions

与语言病理学家在亲子互动中的人机对齐多模态大语言模型

Weiyan Shi, Kenny Tsu Wei Choo

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract)

AI总结 本研究通过多模态大语言模型与语言病理学家的对齐,开发了支持亲子互动分析的系统,实现了85%的感知线索提取准确率和75%的判断精确率,并提出了行为观察-判断系统的构建指南。

Comments This is an earlier version of the work released in May 2025. The version accepted at CHI 2026 is available as a separate preprint at arXiv:2511.04366

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04781 2026-01-09 cs.HC 82%

Dynamic Thermal Feedback in Highly Immersive VR Scenarios: a Multimodal Analysis of User Experience

高度沉浸式VR场景中的动态热反馈:用户体验的多模态分析

Sophie Villenave, Pierre Raimbaud, Guillaume Lavoué

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract)

AI总结 本研究通过多模态分析探讨了高度沉浸式VR场景中动态热反馈对用户体验的影响,发现热反馈可提升沉浸感,但不同质量水平对用户体验的影响因人而异。

Comments 15 pages, 9 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13673 2025-12-30 cs.MM cs.CV cs.IR cs.LG eess.AS 82%

Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

深度音频视觉相关性学习的最新进展与挑战

Luís Vilaça, Yi Yu, Paula Viana

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

AI总结 本文综述了深度音频-视觉相关性学习的最新进展,分析了现有模型、优化方法及未来研究方向。

Comments 8 pages, 1 figure

Journal ref ACM Computing Surveys, 57(12), 1-46, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14115 2025-12-17 cs.SD cs.LG 82%

Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting

多模态对比学习联合框架用于鲁棒的语音术语检测和关键词 spotting

Ramesh Gundluru, Shubham Gupta, Sri Rama Murty K

机构 * Electrical Engineering(电子工程) Artificial Intelligence(人工智能)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出一种联合多模态对比学习框架,通过统一音频和跨模态监督提升语音术语检测和关键词 spotting 的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04720 2025-12-05 cs.SD 82%

M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis

M3-TTS: 多模态DiT对齐与梅尔-潜在空间用于零样本高质量语音合成

Xiaopeng Wang, Chunyu Qiang, Ruibo Fu, Zhengqi Wen, Xuefei Liu, Yukun Liu, Yuzhe Liang, Kang Yin, Yuankun Xie, Heng Xie, Chenxing Li, Chen Zhang, Changsheng Li

机构 * Beijing Institute of Technology(北京理工大学) Kuaishou Technology(快手科技) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 音频语音多模态 :multi-modal(title,abstract);cross-modal(abstract)

AI总结 M3-TTS通过多模态扩散变换器实现零样本高质量语音合成,采用联合扩散层和梅尔-变体编码器,实现高效且自然的语音生成。

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17323 2025-11-24 cs.SD cs.AI cs.CL cs.MM 82%

MusicAIR: A Multimodal AI Music Generation Framework Powered by an Algorithm-Driven Core

MusicAIR: 一种由算法驱动核心的多模态AI音乐生成框架

Callie C. Liao, Duoduo Liao, Ellie L. Zhang

机构 * Stanford University(斯坦福大学) George Mason University(乔治·马歇尔大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 MusicAIR通过算法驱动的核心生成音乐,能从歌词、文本和图像生成符合音乐理论的旋律谱,提升音乐创作效率并降低入门门槛。

Comments Accepted by IEEE Big Data 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01094 2025-11-21 cs.SD cs.AI cs.MM eess.AS 82%

MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions

MMVA: 基于估值和唤醒度的多模态匹配

Suhwan Choi, Kyu Won Kim, Myungjoo Kang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM、eess.AS

AI总结 MMVA通过多模态匹配基于估值和唤醒度,实现跨图像、音乐和音乐字幕的情感内容捕捉,并在估值-唤醒度预测任务中取得最佳性能。

Comments Paper accepted in Artificial Intelligence for Music workshop at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10069 2025-11-12 cs.DC cs.LG 82%

ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism

Zedong Liu, Shenggan Cheng, Guangming Tan, Yang You, Dingwen Tao

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Electronic Science and Technology of China(电子科技大学) National University of Singapore(新加坡国立大学)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract)

Comments Accepted at NeurIPS 2025 Oral (Thirty-Ninth Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19281 2025-11-12 cs.RO 82%

Audio-Visual Traffic Light State Detection for Urban Robots

Sagar Gupta, Akansel Cosgun

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);multi-modal(abstract)

Comments Submitted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2024

Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06988 2025-11-11 cs.LG cs.HC 82%

HCFSLN: Adaptive Hyperbolic Few-Shot Learning for Multimodal Anxiety Detection

Aditya Sneh, Nilesh Kumar Sahu, Anushka Sanjay Shelke, Arya Adyasha, Haroon R. Lone

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06288 2025-11-11 cs.SD cs.CL cs.MM eess.AS 82%

ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction

Wenxuan Wu, Shuai Wang, Xixin Wu, Helen Meng, Haizhou Li

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CL、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13145 2025-11-03 cs.SD 82%

UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model

Yudong Yang, Xiaokang Liu, Shaofeng zhao, Rongfeng Su, Nan Yan, Lan Wang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, China(深圳先进技术研究院,中国科学院,中国) Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences, China(生物医学成像科学与系统重点实验室,中国科学院,中国)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17113 2025-10-28 cs.CV cs.AI cs.CL 82%

MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation

Shoubin Yu, Yue Zhang, Ziyang Wang, Jaehong Yoon, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2025 Findings; The first two authors contributed equally; Github link: https://github.com/Yui010206/MEXA

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02236 2025-10-22 cs.CV cs.MM cs.SD eess.AS 82%

3D Audio-Visual Segmentation

Artem Sokolov, Swapnil Bhosale, Xiatian Zhu

机构 * University of Surrey, UK(Surrey大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at the NeurIPS 2024 Workshop on Audio Imagination; this version updates the project page link

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13182 2025-10-16 cs.LG 82%

Information-Theoretic Criteria for Knowledge Distillation in Multimodal Learning

Rongrong Xie, Yizhou Xu, Guido Sanguinetti

机构 * Scuola Internazionale Superiore di Studi Avanzati (SISSA)(国际先进研究高等学院) École Polytechnique Fédérale de Lausanne (EPFL)(日内瓦联邦理工学院)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06060 2025-10-08 cs.MM cs.AI cs.CV 82%

Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information

Christian Marinoni, Riccardo Fosco Gramaccioni, Eleonora Grassucci, Danilo Comminiello

机构 * Sapienza University of Rome, Italy(罗马大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02165 2025-10-03 cs.SE 82%

Towards fairer public transit: Real-time tensor-based multimodal fare evasion and fraud detection

Peter Wauyo, Dalia Bwiza, Alain Murara, Edwin Mugume, Eric Umuhoza

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01690 2025-10-03 cs.GR cs.HC 82%

Multimodal Feedback for Task Guidance in Augmented Reality

Hu Guo, Lily Patel, Rohan Gupt

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15017 2025-09-30 cs.CL cs.AI cs.SD eess.AS 82%

DM-Codec: Distilling Multimodal Representations for Speech Tokenization

Md Mubtasim Ahasan, Md Fahim, Tasnim Mohiuddin, A K M Mahbubur Rahman, Aman Chadha, Tariq Iqbal, M Ashraful Amin, Md Mofijul Islam, Amin Ahsan Ali

机构 * Center for Computational & Data Sciences, Independent University, Bangladesh(计算与数据科学中心,独立大学,孟加拉国) Amazon GenAI(亚马逊生成人工智能) Qatar Computing Research Institute(卡塔尔计算研究所) University of Virginia(弗吉尼亚大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23833 2025-09-30 eess.AS cs.CV cs.MM cs.SD 82%

AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines

Cancan Li, Fei Su, Juan Liu, Hui Bu, Yulong Wan, Hongbin Suo, Ming Li

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Beijing AISHELL Technology Co., Ltd.(北京AISHELL科技有限公司) AI Center, OPPO(OPPO人工智能中心)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20724 2025-09-26 cs.SI cs.CL cs.CV cs.MM 82%

Visual Authority and the Rhetoric of Health Misinformation: A Multimodal Analysis of Social Media Videos

Mohammad Reza Zarei, Barbara Stead-Coyle, Michael Christensen, Sarah Everts, Majid Komeili

机构 * School of Computer Science(计算机科学学院) Carleton University(卡尔顿大学) Department of Law and Legal Studies(法律与法律研究系) School of Journalism and Communication(新闻与传播学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19631 2025-09-25 eess.AS cs.AI cs.CL 82%

Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning

Shaoshi Ling, Gang Liu, Guoli Ye, Jinyu Li

机构 * Microsoft CoreAI(微软核心人工智能)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18816 2025-09-24 cs.SD cs.CL cs.MM eess.AS 82%

Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models

Junyu Wang, Ziyang Ma, Zhengding Luo, Tianrui Wang, Meng Ge, Xiaobao Wang, Longbiao Wang

机构 * 1Laboratory of Cognitive Computing Application, College of Intelligence Computing, Tianjin University, Tianjin, China 2School of Computer Science, Shanghai Jiao Tong University, Shanghai, China 3School of Electrical \& Electronic Engineering, Nanyang Technological University, Singapore 4Huiyan Technology (Tianjin) Co., Ltd, Tianjin, China

专题命中 音频语音多模态 :cross-modal(title);multi-modal(abstract);分类 cs.CL、cs.MM、eess.AS

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02823 2025-09-01 cs.SD cs.AI cs.MM eess.AS 82%

A Multimodal Symphony: Integrating Taste and Sound through Generative AI

Matteo Spanio, Massimiliano Zampini, Antonio Rodà, Franco Pierucci

机构 * Centro di Sonologia Computazionale (CSC)(计算声学中心) Department of Information Engineering University of Padova(信息工程系帕多瓦大学) Center for Mind/Brain Sciences (CIMeC)(心智/大脑科学中心) University of Trento(特伦托大学) SoundFood s.r.l.(SoundFood公司)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM、eess.AS

Comments 17 pages, 6 figures (2 + 2 figures with 2 subfigures each)

Journal ref Front. Comput. Sci. 7:1575741 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20359 2025-08-29 cs.IR 82%

Progressive Semantic Residual Quantization for Multimodal-Joint Interest Modeling in Music Recommendation

Shijia Wang, Tianpei Ouyang, Qiang Xiao, Dongjing Wang, Yintao Ren, Songpei Xu, Da Guo, Chuanjiang Luo

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05087 2025-08-08 cs.MM cs.AI cs.CL cs.CR 82%

JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering

Renmiao Chen, Shiyao Cui, Xuancheng Huang, Chengwei Pan, Victor Shea-Jay Huang, QingLin Zhang, Xuan Ouyang, Zhexin Zhang, Hongning Wang, Minlie Huang

机构 * CoAI group, DCST, Tsinghua University(清华大学DCST学院) Beihang University(北航大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments 10 pages, 3 tables, 2 figures, to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23010 2025-08-01 cs.LG cs.AI cs.CV cs.SD eess.AS 82%

Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods

Siwoo Park

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06273 2025-07-22 cs.CV cs.MM cs.SD eess.AS 82%

Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

Jeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis, Yong Man Ro

机构 * KAIST(韩国科学技术院) Imperial College London(伦敦帝国理工学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at ICCV 2025. Code available at: https://github.com/JeongHun0716/zero-avsr

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00071 2025-07-01 cs.CV cs.CL cs.MM 82%

I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue

Esam Ghaleb, Bulat Khaertdinov, Aslı Özyürek, Raquel Fernández

机构 * Multimodal Language Department, Max Planck Institute for Psycholinguistics(马克斯·普朗克心理语言学研究所多模态语言部门) Donders Institute for Brain, Cognition and Behaviour, Radboud University(拉德堡德大学大脑、认知与行为研究所) Department of Advanced Computing Sciences, Maastricht University(马斯特里赫特大学高级计算科学系) Institute for Logic, Language and Computation, University of Amsterdam(阿姆斯特丹大学逻辑、语言与计算研究所)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted to Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏