arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-05 至 2026-03-05 共收录 3 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 3 篇

2603.04128 2026-03-05 cs.CV cs.AI cs.MM 85%

Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

Crab$^{+}$: 一种可扩展且统一的音频视觉场景理解模型,具有显式合作

Dongnuan Cai, Henghui Du, Chang Zhou, Xi Chen, Dan Guo, Hongyuan Zhang, Xuelong Li, Di Hu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学耿丽人工智能学院) Institute of Artificial Intelligence of China Telecom (TeleAI)(中国电信人工智能研究院) AI Technology Center, Online Video Business Unit, Tencent PCG(腾讯PCG在线视频业务单元AI技术中心) Hefei University of Technology(合肥工业大学) The University of Hong Kong(香港大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 Crab$^{+}$通过显式合作解决音频视觉任务异质性问题,实现更广泛的任务覆盖和优于单任务模型的性能表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03811 2026-03-05 cs.SD cs.MM eess.AS 84%

Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement

鲁棒的基于大语言模型的音频视觉语音识别与稀疏模态对齐和视觉单元引导的细化

Fei Su, Cancan Li, Juan Liu, Wei Ju, Hongbin Suo, Ming Li

机构 * School of Computer Science, Wuhan University, China(武汉大学计算机学院) School of Artificial Intelligence, Wuhan University, China(武汉大学人工智能学院) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院) AI Center, OPPO, China(OPPO人工智能中心) Digital Innovation Research Center, Duke Kunshan University, China(杜克大学昆山数字创新研究中心)

专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.MM、eess.AS

AI总结 本文提出AVUR-LLM,通过稀疏模态对齐和视觉单元引导细化,提升音频视觉语音识别的鲁棒性,在LRS3数据集上取得显著性能提升。

Comments submitted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08580 2026-03-05 cs.SD cs.AI eess.AS 81%

LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection

LadderSym: 一种用于音乐练习错误检测的多模态交错Transformer

Benjamin Shiue-Hal Chou, Purvish Jajal, Nick John Eliopoulos, James C. Davis, George K. Thiruvathukal, Kristen Yeon-Ji Yun, Yung-Hsiang Lu

机构 * Purdue University(普渡大学) Loyola University Chicago(芝加哥洛约拉大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、eess.AS

AI总结 LadderSym通过双流编码器和多模态策略提升音乐练习错误检测的F1分数,显著提高遗漏和额外音符的识别准确率。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏