arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-05-13 至 2026-05-13 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 6 篇

2605.10815 2026-05-13 cs.AI eess.AS 88%

Probing Cross-modal Information Hubs in Audio-Visual LLMs

探测音频-视觉大语言模型中的跨模态信息枢纽

Jihoo Jung, Chaeyoung Jung, Ji-Hoon Kim, Joon Son Chung

机构 * Department of Electrical Engineering, Korea Advanced Institute of Science The Graduate School of Advanced Imaging Science, Multimedia \& Film, Chung-Ang University, Seoul, Republic of Korea

专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(title,abstract);分类 cs.AI、eess.AS

AI总结 本文研究了音频-视觉大语言模型中音频与视觉模态间的跨模态信息流动,发现信息主要存储在sink tokens中,并提出一种无需训练的hallucination缓解方法。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06371 2026-05-13 cs.CL cs.AI 82%

OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA

OASIS:一个多语言多模态数据集用于文化导向的语音视觉问答

Firoj Alam, Ali Ezzat Shahroor, Md. Arid Hasan, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Mohamed Bayan Kmainasi, Shammur Absar Chowdhury, Basel Mousi, Fahim Dalvi, Nadir Durrani, Natasa Milic-Frayling

机构 * Qatar Computing Research Institute(卡塔尔计算研究 institute)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI;multimodal foundation model(comments)

AI总结 OASIS是一个多语言多模态数据集,涵盖图像、文本和语音,旨在评估模型在现实场景中的文化导向推理能力,包含0.92M真实图像和14.8M问答对。

Comments Multimodal Foundation Models, Large Language Models, Native, Multilingual, Language Diversity, Contextual Understanding, Culturally Informed

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07668 2026-05-13 cs.CV cs.AI cs.LG cs.RO 81%

Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making

观察与聆听内外:多模态人工智能系统用于驾驶员安全评估和智能车辆决策

Ross Greer, Laura Fleig, Maitrayee Keskar, Erika Maquiling, Giovanni Tapia Lopez, Angel Martinez-Sanchez, Parthib Roy, Jake Rattigan, Mira Sur, Alejandra Vidrio, Thomas Marcotte, Mohan Trivedi

机构 * Machine Intelligence, Interaction, and Imagination (Mi3) Laboratory(机器智能、交互与想象实验室) Laboratory for Intelligent and Safe Automobiles (LISA)(智能与安全汽车实验室) Johns Hopkins University(约翰霍普金斯大学) Center for Medicinal Cannabis Research (CMCR)(医药大麻研究中心)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出L-LIO框架,通过融合音频与视觉数据提升驾驶员状态评估和环境理解,探讨音频在车辆安全中的作用,包括驾驶员语音分类、乘客指令分析及视觉不足时的音频辅助。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11605 2026-05-13 cs.CV cs.AI 79%

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs

保留音频无法表达的内容:面向多模态大语言模型的上下文保留令牌修剪

Chaeyoung Jung, Kyeongha Rho, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI

AI总结 本文提出ContextGuard框架,通过保留音频-视觉上下文并去除跨模态冗余,实现多模态大语言模型的高效令牌修剪,实验表明其在多个基准测试中表现优异,能有效减少输入令牌数量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11700 2026-05-13 cs.HC 78%

MindMirror: A Local-First Multimodal State-Aware Support System for Digital Workers

MindMirror: 一种面向数字工作者的本地优先多模态状态感知支持系统

Wenqi Luo, Changbo Wang, Yan Wang

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 MindMirror通过整合面部表情、文本输入和语音交互等多模态数据,为数字工作者提供本地优先的状态感知支持,通过本地大语言模型生成响应,提升工作效率与状态监控能力。

Comments 10 pages, 4 figures, 12 tables. Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12002 2026-05-13 cs.LG 50%

Detecting In-Person Conversations in Noisy Real-World Environments with Smartwatch Audio and Motion Sensing

利用智能手表音频和运动传感检测嘈杂真实环境中的面对面对话

Alice Zhang, Callihan Bertley, Dawei Liang, Edison Thomaz

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本文提出一种新型方法,通过智能手表融合音频和运动数据,检测面对面对话,实验表明融合惯性数据能显著提升检测性能,实验室和半自然环境下分别达到82.0%和77.2%的宏F1分数。

Comments Accepted to ACM Transactions on Intelligent Systems and Technology

详情

展开后加载摘要…

URL PDF HTML 收藏