arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4562 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4562 篇

2511.09989 2025-11-14 cs.LG 78%

Towards Robust Multimodal Learning in the Open World

Fushuo Huo

机构 * Department of Computing(计算系)

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09039 2025-11-13 cs.LG cs.CY cs.HC 78%

Fairness-Aware Few-Shot Learning for Audio-Visual Stress Detection

Anushka Sanjay Shelke, Aditya Sneh, Arya Adyasha, Haroon R. Lone

专题命中 音频语音多模态 :audio-visual(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05304 2025-11-10 cs.HC 78%

psiUnity: A Platform for Multimodal Data-Driven XR

Akhil Ajikumar, Sahil Mayenkar, Steven Yoo, Sakib Reza, Mohsen Moghaddam

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26830 2025-11-03 cs.LG cs.CR 78%

SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation

Guangzhi Su, Shuchang Huang, Yutong Ke, Zhuohang Liu, Long Qian, Kaizhu Huang

机构 * Independent Researcher(独立研究者)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10396 2025-10-20 cs.SD 78%

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations

Wenxiang Guo, Changhao Pan, Zhiyuan Zhu, Xintong Hu, Yu Zhang, Li Tang, Rui Yang, Han Wang, Zongbao Zhang, Yuhan Wang, Yixuan Chen, Hankun Xu, Ke Xu, Pengfei Fan, Zhetao Chen, Yanhao Yu, Qiange Huang, Fei Wu, Zhou Zhao

机构 * Zhejiang University(浙江大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11596 2025-10-14 cs.HC 78%

GlobalizeEd: A Multimodal Translation System that Preserves Speaker Identity in Academic Lectures

Hoang-Son Vo, Karina Kolmogortseva, Ngumimi Karen Iyortsuun, Hong-Duyen Vo, Soo-Hyung Kim

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22425 2025-10-13 cs.SD 78%

From Coarse to Fine: Recursive Audio-Visual Semantic Enhancement for Speech Separation

Ke Xue, Rongfei Fan, Lixin, Dawei Zhao, Chao Zhu, Han Hu

机构 * School of Cyberspace Science and Technology(网络空间科学与技术学院) Beijing Institute of Technology(北京理工大学) Sun Yat-sen University(中山大学) Qilu University of Technology(齐鲁工业大学) Shandong Computer Science Center(山东计算机科学中心) School of Information and Electronics(信息电子学院)

专题命中 音频语音多模态 :audio-visual(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07839 2025-10-09 cs.RO 78%

Touch Speaks, Sound Feels: A Multimodal Approach to Affective and Social Touch from Robots to Humans

Qiaoqiao Ren, Tony Belpaeme

机构 * Faculty of Engineering and Architecture, IDLab-AIRO, Ghent University – imec, Technologiepark 126, 9052 Gent, Belgium(工程与建筑学院,IDLab-AIRO,根特大学–imec,Technologiepark 126,9052 Gent,比利时)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06872 2025-10-09 cs.HC 78%

Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz

Frederic Gmeiner, Kenneth Holstein, Nikolas Martelaro

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 18 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19600 2025-09-25 cs.HC 78%

vashTimer: A Multi-Purpose, Multimodal Mobile App For Maintaining Passage of Time by means of Visual, Auditory, Speech, and Haptic Alerts

Aziz N Zeidieh, Sanchita S. Kamath, JooYoung Seo

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18706 2025-09-24 cs.HC 78%

M4SER: Multimodal, Multirepresentation, Multitask, and Multistrategy Learning for Speech Emotion Recognition

Jiajun He, Xiaohan Shi, Cheng-Hung Hu, Jinyi Mi, Xingfeng Li, Tomoki Toda

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted by IEEE Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16920 2025-09-23 cs.RO cs.HC 78%

SwarmChat: An LLM-Based, Context-Aware Multimodal Interaction System for Robotic Swarms

Ettilla Mohiuddin Eumi, Hussein Abbass, Nadine Marcus

机构 * School of Systems & Computing, UNSW Canberra(系统与计算学院,UNSW堪培拉) School of Computer Science and Engineering, UNSW Sydney(计算机科学与工程学院,UNSW悉尼)

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This paper has been accepted and presented at the 16th International Conference on Swarm Intelligence (ICSI 2025), held on July 11-15, 2025, in Yokohama, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02359 2025-09-12 cs.LG 78%

Attribution Regularization for Multimodal Paradigms

Sahiti Yerramilli, Jayant Sravan Tamarapalli, Jonathan Francis, Eric Nyberg

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08689 2025-09-11 cs.HC 78%

Augmenting speech transcripts of VR recordings with gaze, pointing, and visual context for multimodal coreference resolution

Riccardo Bovo, Frederik Brudy, George Fitzmaurice, Fraser Anderson

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06382 2025-09-09 cs.HC 78%

Context-Adaptive Hearing Aid Fitting Advisor through Multi-turn Multimodal LLM Conversation

Yingke Ding, Zeyu Wang, Xiyuxing Zhang, Hongbin Chen, Zhenan Xu

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Ubicomp Companion 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15632 2025-08-26 cs.SD 78%

ASCMamba: Multimodal Time-Frequency Mamba for Acoustic Scene Classification

Bochao Sun, Dong Wang, ZhanLong Yang, Jun Yang, Han Yin

机构 * School of Marine Science and Technology, Northwestern Polytechnical University, Xi’an, China(海洋科学与技术学院,西北工业大学,西安,中国) School of Automation, Northwestern Polytechnical University, Xi’an, China(自动化学院,西北工业大学,西安,中国) School of Electrical Engineering, KAIST, Daejeon, Republic of Korea(电气工程学院,韩国成均馆大学,大田,韩国)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16320 2025-08-25 physics.ed-ph 78%

AI-Supported Mini-Labs: Combining Smartphone-Based Experiments and Multimodal AI

Jochen Kuhn, David J. Rakestraw, Stefan Küchemann, Patrik Vogt

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15565 2025-08-22 cs.SD 78%

Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization

Liping Chen, Chenyang Guo, Rui Wang, Kong Aik Lee, Zhenhua Ling

机构 * University of Science and Technology of China(中国科学技术大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 音频语音多模态 :any-to-any(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14976 2025-08-22 cs.LG 78%

Aura-CAPTCHA: A Reinforcement Learning and GAN-Enhanced Multi-Modal CAPTCHA System

Joydeep Chandra, Prabal Manhas, Ramanjot Kaur, Rashi Sahay

机构 * Department of Computer Science and Engineering, Chandigarh University, Mohali, Punjab, India(昌迪加尔大学计算机科学与工程系,莫哈利,旁遮普,印度) Department of Computer Science and Engineering, Manav Rachna International Institute of Research and Studies(曼纳瓦拉国际研究与学习研究所计算机科学与工程系)

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14912 2025-08-22 cs.IR 78%

Multimodal Recommendation via Self-Corrective Preference Alignmen

Yalong Guan, Xiang Chen, Mingyang Wang, Xiangyu Wu, Lihao Liu, Chao Qi, Shuang Yang, Tingting Gao, Guorui Zhou, Changjian Chen

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14631 2025-08-21 cs.SE 78%

Towards a DSL to Formalize Multimodal Requirements

Marcos Gomez-Vazquez, Jordi Cabot

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21154 2025-08-21 cs.HC 78%

Hypergraph Multi-Modal Learning for EEG-based Emotion Recognition in Conversation

Zijian Kang, Yueyang Li, Shengyu Gong, Weiming Zeng, Hongjie Yan, Lingbin Bian, Zhiguo Zhang, Wai Ting Siok, Nizhuan Wang

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09402 2025-08-14 cs.HC 78%

Realtime Multimodal Emotion Estimation using Behavioral and Neurophysiological Data

Von Ralph Dane Marquez Herbuela, Yukie Nagai

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03300 2025-08-13 cs.RO cs.LG 78%

Touch and Tell: Multimodal Decoding of Human Emotions and Social Gestures for Robots

Qiaoqiao Ren, Remko Proesmans, Yuanbo Hou, Francis wyffels, Tony Belpaeme

机构 * Faculty of Engineering and Architecture(工程与建筑学院) IDLab-AIRO, Ghent University – imec(IDLab-AIRO,根特大学–imec) Department of Engineering Science, University of Oxford(工程科学系,牛津大学)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06117 2025-08-11 cs.HC 78%

A Multimodal Framework for Understanding Collaborative Design Processes

Maurice Koch, Nelusa Pathmanathan, Daniel Weiskopf, Kuno Kurzhals

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted to IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04906 2025-07-31 cs.MM cs.CV cs.SD eess.AS 78%

Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation

Ivan Rinaldi, Nicola Fanelli, Giovanna Castellano, Gennaro Vessio

机构 * Department of Computer Science, University of Bari Aldo Moro, Italy(巴里阿尔多·莫罗大学计算机科学系)

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CV、cs.MM、eess.AS

Comments Presented at the AI for Visual Arts (AI4VA) workshop at ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15826 2025-07-22 cs.IR cs.LG 78%

Just Ask for Music (JAM): Multimodal and Personalized Natural Language Music Recommendation

Alessandro B. Melchiorre, Elena V. Epure, Shahed Masoudian, Gustavo Escobedo, Anna Hausberger, Manuel Moussallam, Markus Schedl

机构 * Johannes Kepler University Linz(约翰· Kepler大学林茨) Criteo AI Lab(Criteo人工智能实验室) Deezer Research(Deezer研究) Johannes Kepler University Linz and Linz Institute of Technology(约翰· Kepler大学林茨和林茨技术学院)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01973 2025-07-15 cs.CE 78%

Multimodal Financial Foundation Models (MFFMs): Progress, Prospects, and Challenges

Xiao-Yang Liu Yanglet, Yupeng Cao, Li Deng

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14445 2025-06-18 cs.IR 78%

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

Ruofan Hu, Yan Xia, Minjie Hong, Jieming Zhu, Bo Chen, Xiaoda Yang, Minghui Fang, Tao Jin

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted by Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07598 2025-06-16 cs.HC 78%

Towards spatial computing: recent advances in multimodal natural interaction for XR headsets

Zhimin Wang, Maohang Rao, Shanghua Ye, Weitao Song, Feng Lu

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 28 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏