arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4557 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4557 篇

2408.02978 2025-06-25 cs.MM cs.AI cs.CV 82%

ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval

Ruixiang Zhao, Jian Jia, Yan Li, Xuehan Bai, Quan Chen, Han Li, Peng Jiang, Xirong Li

机构 * Renmin University of China(中国人民大学) Kuaishou Technology(快手科技)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments accepted for publication as a REGULAR paper in the IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11064 2025-06-16 eess.AS cs.AI cs.CL cs.SD 82%

PMF-CEC: Phoneme-augmented Multimodal Fusion for Context-aware ASR Error Correction with Error-specific Selective Decoding

Jiajun He, Tomoki Toda

机构 * Graduate School of Informatics, Nagoya University(信息学研究科,名古屋大学) Information Technology Center, Nagoya University(信息科技中心,名古屋大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments Accepted by IEEE TASLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24059 2025-06-02 cs.LG 82%

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition

Sean Foley, Hong Nguyen, Jihwan Lee, Sudarsana Reddy Kadiri, Dani Byrd, Louis Goldstein, Shrikanth Narayanan

机构 * University of Southern CaliforniaUSA(美国南加州大学) Signal Analysis and Interpretation Laboratory(信号分析与解释实验室) Department of Linguistics(语言学系)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12665 2025-05-20 cs.RO 82%

Audio-Visual Contact Classification for Tree Structures in Agriculture

Ryan Spears, Moonyoung Lee, George Kantor, Oliver Kroemer

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10176 2025-05-16 cs.NE 82%

Incorporating brain-inspired mechanisms for multimodal learning in artificial intelligence

Xiang He, Dongcheng Zhao, Yang Li, Qingqun Kong, Xin Yang, Yi Zeng

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract)

Comments The manuscript is under review and the code is available at https://github.com/Brain-Cog-Lab/IEMF

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15171 2025-04-22 cs.LG 82%

Audio-Visual Class-Incremental Learning for Fish Feeding intensity Assessment in Aquaculture

Meng Cui, Xianghu Yue, Xinyuan Qian, Jinzheng Zhao, Haohe Liu, Xubo Liu, Daoliang Li, Wenwu Wang

机构 * Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(视觉、语音和信号处理中心(CVSSP), Surrey大学) College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) School of Computer and Communication Engineering, University of Science and Technology Beijing(计算机与通信工程学院,北京科技大学) National Innovation Center for Digital Fishery, China Agricultural University(数字渔业国家创新中心,中国农业大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18368 2025-04-18 cs.CL cs.AI cs.LG eess.AS 82%

AMPS: ASR with Multimodal Paraphrase Supervision

Abhishek Gupta, Amruta Parulekar, Sameep Chattopadhyay, Preethi Jyothi

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12225 2025-04-10 cs.LG cs.AI cs.CL cs.MM 82%

DLF: Disentangled-Language-Focused Multimodal Sentiment Analysis

Pan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen, Jingtong Hu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments AAAI 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06789 2025-04-07 cs.RO 82%

AV-PedAware: Self-Supervised Audio-Visual Fusion for Dynamic Pedestrian Awareness

Yizhuo Yang, Shenghai Yuan, Muqing Cao, Jianfei Yang, Lihua Xie

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract)

Comments This work has been accepted for publication at the 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). Personal use is permitted. For other uses, permission from IEEE is required

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17698 2025-03-18 cs.CV cs.MM cs.SD eess.AS 82%

Video-Guided Foley Sound Generation with Multimodal Controls

Ziyang Chen, Prem Seetharaman, Bryan Russell, Oriol Nieto, David Bourgin, Andrew Owens, Justin Salamon

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at CVPR 2025. Project site: https://ificl.github.io/MultiFoley/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09291 2025-03-18 cs.MM cs.AI cs.SD eess.AS 82%

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Kyeongha Rho, Hyeongkeun Lee, Valentio Iverson, Joon Son Chung

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI、cs.MM、eess.AS

Comments 5 pages, 2 figures; Accepted to ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05782 2025-03-18 cs.SD cs.CV cs.LG cs.MM eess.AS 82%

Sequential Contrastive Audio-Visual Learning

Ioannis Tsiamas, Santiago Pascual, Chunghsin Yeh, Joan Serrà

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments ICASSP 2025. Version 1 contains more details

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08137 2025-03-14 cs.CV cs.CR cs.MM cs.SD eess.AS 82%

Audio-Visual Deepfake Detection With Local Temporal Inconsistencies

Marcella Astrid, Enjie Ghorbel, Djamila Aouada

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted in ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00358 2025-02-24 cs.SD cs.AI cs.LG cs.MM eess.AS 82%

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

Jia Li, Wenjie Zhao, Ziru Huang, Yunhui Guo, Yapeng Tian

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05307 2025-02-11 cs.CE cs.LG 82%

Audio-visual cross-modality knowledge transfer for machine learning-based in-situ monitoring in laser additive manufacturing

Jiarui Xie, Mutahar Safdar, Lequn Chen, Seung Ki Moon, Yaoyao Fiona Zhao

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract)

Comments 47 pages, 19 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02345 2025-02-11 cs.CV cs.AI cs.LG cs.MM 82%

Progressive Confident Masking Attention Network for Audio-Visual Segmentation

Yuxuan Wang, Jinchao Zhu, Feng Dong, Shuyue Zhu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments 23 pages, 11 figures, submitted to Elsevier Knowledge-Based System

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18157 2025-01-31 cs.SD cs.CV cs.MM eess.AS 82%

Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment

Joanna Hong, Sanjeel Parekh, Honglie Chen, Jacob Donley, Ke Tan, Buye Xu, Anurag Kumar

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16295 2025-01-28 cs.LG cs.AI cs.CL cs.CV 82%

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity

Weixin Liang, Junhong Shen, Genghan Zhang, Ning Dong, Luke Zettlemoyer, Lili Yu

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11409 2025-01-16 cs.CV cs.AI cs.MM 82%

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

Rui Liu, Shuwei He, Yifan Hu, Haizhou Li

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments 9 pages,2 figures, Accepted by AAAI'2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05301 2025-01-16 cs.SD cs.AI cs.CV cs.LG eess.AS eess.SP 82%

Diffusion-based Unsupervised Audio-visual Speech Enhancement

Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel, Xavier Alameda-Pineda

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.AI、eess.AS

Journal ref International Conference on Acoustics Speech and Signal Processing (ICASSP), IEEE, Apr 2025, Hyderabad, India

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06282 2025-01-14 cs.CL cs.AI cs.HC cs.SD eess.AS 82%

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Qian Chen, Yafeng Chen, Yanni Chen, Mengzhe Chen, Yingda Chen, Chong Deng, Zhihao Du, Ruize Gao, Changfeng Gao, Zhifu Gao, Yabin Li, Xiang Lv, Jiaqing Liu, Haoneng Luo, Bin Ma, Chongjia Ni, Xian Shi, Jialong Tang, Hui Wang, Hao Wang, Wen Wang, Yuxuan Wang, Yunlan Xu, Fan Yu, Zhijie Yan, Yexin Yang, Baosong Yang, Xian Yang, Guanrou Yang, Tianyu Zhao, Qinglin Zhang, Shiliang Zhang, Nan Zhao, Pei Zhang, Chong Zhang, Jinren Zhou

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments Work in progress. Authors are listed in alphabetical order by family name

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16928 2024-12-24 cs.SD cs.CV cs.MM eess.AS 82%

AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification

Zhenyuan Xiao, Yizhuo Yang, Guili Xu, Xianglong Zeng, Shenghai Yuan

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Submitted to ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15995 2024-12-23 cs.CL cs.AI cs.SD eess.AS 82%

Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling

Maximillian Chen, Ruoxi Sun, Sercan Ö. Arık

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments 22 pages, 6 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11715 2024-12-17 cs.CV cs.MM cs.SD eess.AS 82%

Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning

RunLin Yu, Yipu Gong, Wenrui Li, Aiwen Sun, Mengren Zheng

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10460 2024-12-17 cs.CV cs.AI cs.SD eess.AS 82%

Enriching Multimodal Sentiment Analysis through Textual Emotional Descriptions of Visual-Audio Content

Sheng Wu, Xiaobao Wang, Longbiao Wang, Dongxiao He, Jianwu Dang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、eess.AS

Journal ref AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08161 2024-12-12 cs.CV cs.LG cs.MM cs.SD eess.AS 82%

Collaborative Hybrid Propagator for Temporal Misalignment in Audio-Visual Segmentation

Kexin Li, Zongxin Yang, Yi Yang, Jun Xiao

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02142 2024-12-04 cs.CV cs.AI cs.CL cs.IR 82%

Personalized Multimodal Large Language Models: A Survey

Junda Wu, Hanjia Lyu, Yu Xia, Zhehao Zhang, Joe Barrow, Ishita Kumar, Mehrnoosh Mirtaheri, Hongjie Chen, Ryan A. Rossi, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang, Xiang Chen, Hanieh Deilamsalehy, Namyong Park, Sungchul Kim, Huanrui Yang, Subrata Mitra, Zhengmian Hu, Nedim Lipka, Dang Nguyen, Yue Zhao, Jiebo Luo, Julian McAuley

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12306 2024-11-13 cs.CL cs.CV cs.SD eess.AS 82%

Measuring Sound Symbolism in Audio-visual Models

Wei-Cheng Tseng, Yi-Jen Shih, David Harwath, Raymond Mooney

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.CL、eess.AS

Comments SLT 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02551 2024-11-08 cs.SD cs.AI cs.MM eess.AS 82%

PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text

Hayeon Bang, Eunjin Choi, Megan Finch, Seungheon Doh, Seolhee Lee, Gyeong-Hoon Lee, Juhan Nam

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 cs.AI、cs.MM、eess.AS

Comments Accepted for publication at the 3rd Workshop on NLP for Music and Audio (NLP4MusA 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00562 2024-11-05 eess.AS cs.CV cs.MM cs.SD 82%

Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification

Aref Farhadipour, Masoumeh Chapariniya, Teodora Vukovic, Volker Dellwo

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments This paper was accepted at the ICNLSP2024 conference

详情

展开后加载摘要…

URL PDF HTML 收藏