arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4559 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4559 篇

2506.23623 2025-07-01 cs.CV 79%

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

Shaofei Huang, Rui Ling, Tianrui Hui, Hongyu Li, Xu Zhou, Shifeng Zhang, Si Liu, Richang Hong, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) Chinese Academy of Sciences(中国科学院) Beihang University(北航) Sangfor Technologies(深信服技术)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025; Code: https://github.com/spyflying/VCT_AVS; Models: https://huggingface.co/nowherespyfly/VCT_AVS

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23271 2025-07-01 cs.CV 79%

Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation

Jinxing Zhou, Zhihui Li, Yongqiang Yu, Yanghao Zhou, Ruohao Guo, Guangyao Li, Yuxin Mao, Mingfei Han, Xiaojun Chang, Meng Wang

机构 * MBZUAI Hefei University of Technology(合肥工业大学) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Peking University(北京大学) Tsinghua University(清华大学) OpenNLP Lab(OpenNLP实验室)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22926 2025-07-01 cs.HC cs.GR cs.MM 79%

Coordinated 2D-3D Visualization of Volumetric Medical Data in XR with Multimodal Interactions

Qixuan Liu, Shi Qiu, Yinqiao Wang, Xiwen Wu, Kenneth Siu Ho Chok, Chi-Wing Fu, Pheng-Ann Heng

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments IEEE VIS 2025 Short Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.24066 2025-07-01 eess.AS eess.SP 79%

Cough-E: A multimodal, privacy-preserving cough detection algorithm for the edge

Stefano Albini, Lara Orlandic, Jonathan Dan, Jérôme Thevenot, Tomas Teijeiro, Denisa Andreea Constantinescu, David Atienza

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20945 2025-06-27 cs.SD eess.AS 79%

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis

Rui Niu, Weihao Wu, Jie Chen, Long Ma, Zhiyong Wu

机构 * Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(深圳国际研究生院,清华大学,深圳,中国)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19603 2025-06-25 cs.CL cs.SI 79%

Social Hatred: Efficient Multimodal Detection of Hatemongers

Tom Marzea, Abraham Israeli, Oren Tsur

机构 * Ben Gurion University(本古里安大学) University of Michigan(密歇根大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments To be published in WOAH, July 2025. arXiv admin note: text overlap with arXiv:2409.14464

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13419 2025-06-17 eess.IV cs.CV 79%

Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos

Riku Takahashi, Ryugo Morita, Jinjia Zhou

机构 * Hosei University(恒生大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted to ICMR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10331 2025-06-13 cs.CV eess.IV 79%

Research on Audio-Visual Quality Assessment Dataset and Method for User-Generated Omnidirectional Video

Fei Zhao, Da Pan, Zelu Qi, Ping Shi

机构 * School of Information and Communication Engineering(信息与通信工程学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Our paper has been accepted by ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06759 2025-06-10 cs.CV 79%

LitMAS: A Lightweight and Generalized Multi-Modal Anti-Spoofing Framework for Biometric Security

Nidheesh Gorthi, Kartik Thakral, Rishabh Ranjan, Richa Singh, Mayank Vatsa

机构 * Indian Institute of Information Technology Kottayam(印度信息技术学院科塔亚姆) Indian Institute of Technology Jodhpur(印度理工学院朱罗普尔)

专题命中 音频语音多模态 :multi-modal(title);cross-modal(abstract);分类 cs.CV

Comments Accepted in Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03980 2025-06-05 cs.CL 79%

Voice Activity Projection Model with Multimodal Encoders

Takeshi Saga, Catherine Pelachaud

机构 * Sorbonne University(索邦大学) CNRS(国家科学研究中心)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02470 2025-06-04 cs.AI 79%

A Smart Multimodal Healthcare Copilot with Powerful LLM Reasoning

Xuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan Miao

机构 * Joint NTU-UBC Research Centre of Excellence in Active Living for the Elderly (LILY), NTU(联合NTU-UBC老龄化积极生活卓越研究中心(LILY),NTU) College of Computing and Data Science, Nanyang Technological University (NTU), Singapore(计算与数据科学学院,南洋理工大学(NTU),新加坡) Tan Tock Seng Hospital, Singapore(坦 tok sing 医院,新加坡) Woodlands Health, Singapore(伍德兰兹健康,新加坡)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02178 2025-06-04 cs.SD cs.CL 79%

Cocktail-Party Audio-Visual Speech Recognition

Thai-Binh Nguyen, Ngoc-Quan Pham, Alexander Waibel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CL

Comments Accepted at Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01270 2025-06-03 eess.AS cs.SD 79%

Online Audio-Visual Autoregressive Speaker Extraction

Zexu Pan, Wupeng Wang, Shengkui Zhao, Chong Zhang, Kun Zhou, Yukun Ma, Bin Ma

机构 * Alibaba Group(阿里巴巴集团) Singapore National University of Singapore(新加坡国立大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Interspeech2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07217 2025-06-03 cs.SD cs.CV 79%

ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation

Zixuan Wang, Chi-Keung Tang, Yu-Wing Tai

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Dartmouth College(达特茅斯学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments Project page: https://vincent2311.github.io/ReelWave_demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23298 2025-05-30 cs.SD cs.IR eess.AS 79%

Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning

Xiaofeng Pan, Jing Chen, Haitong Zhang, Menglin Xing, Jiayi Wei, Xuefeng Mu, Zhongqian Xie

机构 * NetEase Inc.(网易公司)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments ICMR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10034 2025-05-30 cs.AI 79%

The First MPDD Challenge: Multimodal Personality-aware Depression Detection

Changzeng Fu, Zelin Fu, Qi Zhang, Xinhe Kuang, Jiacheng Dong, Kaifeng Su, Yikai Su, Wenbo Shi, Junfeng Yao, Yuliang Zhao, Shiqi Zhao, Jiadong Wang, Siyang Song, Chaoran Liu, Yuichiro Yoshikawa, Björn Schuller, Hiroshi Ishiguro

机构 * Northeastern University(东北大学) University of Technology Sydney(悉尼大学) Xiamen University(厦门大学) Technical University of Munich(慕尼黑技术大学) University of Cambridge(剑桥大学) National Information Institute(国家信息研究所) Osaka University(大阪大学) Imperial College London(伦敦帝国理工学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments This paper has been accepted as part of the MPDD Challenge in the ACMMM 2025 Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19938 2025-05-27 cs.CV 79%

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning

Wenrui Li, Penghong Wang, Xingtao Wang, Wangmeng Zuo, Xiaopeng Fan, Yonghong Tian

机构 * Harbin Institute of Technology(哈尔滨工业大学) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院) Peking University(北京大学) School of AI for Science(科学人工智能学院) Peng Cheng Laboratory(鹏城实验室)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by IEEE TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11450 2025-05-27 cs.CV 79%

VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition

Puneet Kumar, Sarthak Malik, Balasubramanian Raman, Xiaobai Li

机构 * Center for Machine Vision and Signal Analysis, University of Oulu(机器视觉与信号分析中心,奥卢大学) Indian Institute of Technology Roorkee(印度理工学院罗尔基分校) State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05078 2025-05-26 cs.LG cs.AI cs.IT math.IT 79%

Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data

David Heurtel-Depeiges, Anian Ruoss, Joel Veness, Tim Genewein

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00735 2025-05-20 cs.CR cs.AI cs.SE 79%

`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs

Chun Wai Chiu, Linghan Huang, Bo Li, Huaming Chen, Kim-Kwang Raymond Choo

机构 * School of Electrical and Computer Engineering, The University of Sydney(悉尼大学电气与计算机工程学院) University of Chicago(芝加哥大学) University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22076 2025-05-20 cs.SD cs.HC eess.AS 79%

USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis

Luca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith C. H. Ngai, Chenshu Wu

机构 * The University of Hong Kong(香港大学)

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 eess.AS

Comments Accepted by Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (ACM IMWUT/UbiComp 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07901 2025-05-13 cs.MM 79%

Bridging Discrete and Continuous: A Multimodal Strategy for Complex Emotion Detection

Jiehui Jia, Huan Zhang, Jinhua Liang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01578 2025-05-06 cs.CV 79%

Grounding Task Assistance with Multimodal Cues from a Single Demonstration

Gabriel Sarch, Balasaravanan Thoravi Kumaravel, Sahithya Ravi, Vibhav Vineet, Andrew D. Wilson

机构 * Microsoft Research(微软研究院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21366 2025-05-01 cs.SD cs.AI 79%

DGFNet: End-to-End Audio-Visual Source Separation Based on Dynamic Gating Fusion

Yinfeng Yu, Shiyu Sun

机构 * School of Computer Science and Technology(计算机科学与技术学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI

Comments Main paper (9 pages). Accepted for publication by ICMR(International Conference on Multimedia Retrieval) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17129 2025-04-29 eess.AS cs.SD 79%

Uncovering the Visual Contribution in Audio-Visual Speech Recognition

Zhaofeng Lin, Naomi Harte

机构 * Sigmedia Group, School of Engineering Trinity College Dublin(Sigmedia集团,工程学院,三一学院都柏林)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments 5 pages, 2 figures. Accepted to ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09154 2025-04-15 cs.MM cs.LG 79%

Exploring Modality Disruption in Multimodal Fake News Detection

Moyang Liu, Kaiying Yan, Yukun Liu, Ruibo Fu, Zhengqi Wen, Xuefei Liu, Chenxing Li

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05746 2025-04-09 cs.CV 79%

Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation

Zhihua Xu, Tianshui Chen, Zhijing Yang, Siyuan Peng, Keze Wang, Liang Lin

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted at TMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02988 2025-04-07 cs.SD eess.AS 79%

Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection

Adrian S. Roman, Aiden Chang, Gerardo Meza, Iran R. Roman

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07751 2025-04-03 cs.SD cs.AI cs.CV cs.MM eess.AS 79%

SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model

Xinyuan Qian, Jiaran Gao, Yaodan Zhang, Qiquan Zhang, Hexin Liu, Leibny Paola Garcia, Haizhou Li

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.AI、cs.MM

Comments accepted by IEEE Journal of Selected Topics in Signal Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16376 2025-04-03 cs.AI cs.HC 79%

Beyond Text-to-Text: An Overview of Multimodal and Generative Artificial Intelligence for Education Using Topic Modeling

Ville Heilala, Roberto Araya, Raija Hämäläinen

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref Proceedings of the 40th ACM/SIGAPP Symposium on Applied Computing (SAC'25), March 31--April 4, 2025, Catania, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏