arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4562 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4562 篇

2506.03378 2025-06-05 eess.AS cs.CV cs.MM 78%

SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer

Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Abu Osama Siddiqui, Sarthak Jain, Priyabrata Mallick, Jaya Sai Kiran Patibandla, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma

机构 * IIIT-DelhiIndia(印度德里印度理工学院) V.B.S.P.UIndia(印度V.B.S.P.U) UPESIndia(印度UPES) Independent ResearcherIndia(印度独立研究者) Reliance AIIndia(印度Reliance AI) University of TartuEstonia(爱沙尼亚塔尔图大学) Plaksha UniversityIndia(印度Plaksha大学)

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.MM、eess.AS

Comments Accepted to INTERSPEECH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13338 2025-06-04 cs.CL cs.AI eess.AS 78%

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

Qiongqiong Wang, Hardik B. Sailor, Tianchi Liu, Ai Ti Aw

机构 * Agency for Science, Technology and Research (A ⋆ ⋆ \star ⋆ STAR)(科技研究局) Institute for Infocomm Research (I 2 R)(信息通信研究所)

专题命中 音频语音多模态 :multi-modal(title);分类 cs.CL、cs.AI、eess.AS

Comments Accepted at Interspeech 2025. [v2]: The dataset has been released, and the link is now updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19877 2025-05-05 cs.HC 78%

Towards Multimodal Large-Language Models for Parent-Child Interaction: A Focus on Joint Attention

Weiyan Shi, Viet Hai Le, Kenny Tsu Wei Choo

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted at CHI 2025 Late Breaking Work

Journal ref Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2025, Article No. 535, Pages 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18914 2025-04-29 cs.LG stat.AP stat.ML 78%

Factor Analysis with Correlated Topic Model for Multi-Modal Data

Małgorzata Łazęcka, Ewa Szczurek

机构 * Faculty of Mathematics, Informatics and Mechanics, University of Warsaw(华沙大学数学、信息学与力学学院) Institute of Computer Science, Polish Academy of Sciences(波兰科学院计算机科学研究所) Institute of AI for Health, Helmholtz Center Munich(慕尼黑海德堡中心人工智能与健康研究所)

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract)

Comments AISTATS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14214 2025-04-22 cs.IR 78%

Teach Me How to Denoise: A Universal Framework for Denoising Multi-modal Recommender Systems via Guided Calibration

Hongji Li, Hanwen Du, Youhua Li, Junchen Fu, Chunxiao Li, Ziyi Zhuang, Jiakang Li, Yongxin Ni

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments Accepted to ACM Web Search and Data Mining (WSDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11460 2025-04-21 cs.CV cs.AI cs.CL 78%

Semantic Matters: Multimodal Features for Affective Analysis

Tobias Hallmen, Robin-Nico Kampa, Fabian Deuser, Norbert Oswald, Elisabeth André

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06593 2025-04-10 cs.RO cs.HC 78%

A Multi-Modal Interaction Framework for Efficient Human-Robot Collaborative Shelf Picking

Abhinav Pathak, Kalaichelvi Venkatesan, Tarek Taha, Rajkumar Muthusamy

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00785 2025-04-07 cs.RO 78%

Natural Multimodal Fusion-Based Human-Robot Interaction: Application With Voice and Deictic Posture via Large Language Model

Yuzhi Lai, Shenghai Yuan, Youssef Nassar, Mingyu Fan, Atmaraaj Gopal, Arihiro Yorita, Naoyuki Kubota, Matthias Rätsch

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract)

Comments Accepted for publication by IEEE Robotics & Automation Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19334 2025-03-26 cs.HC 78%

Design of Seamless Multi-modal Interaction Framework for Intelligent Virtual Agents in Wearable Mixed Reality Environment

Ghazanfar Ali, Hong-Quan Le, Junho Kim, Seoung-won Hwang, Jae-In Hwang

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract)

Comments 6 pages, 14 Figures, Computer Animation and Social Agents (CASA 2019)

Journal ref CASA 2019: Proceedings of the 32nd International Conference on Computer Animation and Social Agents - Year 2019 - Pages 47 - 52

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01226 2025-03-04 q-bio.NC cs.LG 78%

Dementia Insights: A Context-Based MultiModal Approach

Sahar Sinene Mehdoui, Abdelhamid Bouzid, Daniel Sierra-Sosa, Adel Elmaghraby

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04592 2025-03-04 cs.HC 78%

CardioAI: A Multimodal AI-based System to Support Symptom Monitoring and Risk Detection of Cancer Treatment-Induced Cardiotoxicity

Siyi Wu, Weidan Cao, Shihan Fu, Bingsheng Yao, Ziqi Yang, Changchang Yin, Varun Mishra, Daniel Addison, Ping Zhang, Dakuo Wang

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17971 2025-02-26 cs.RO cs.HC 78%

Multimodal Interaction and Intention Communication for Industrial Robots

Tim Schreiter, Andrey Rudenko, Jens V. Rüppel, Martin Magnusson, Achim J. Lilienthal

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted to the 1st German Robotics Conference (GRC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16703 2025-02-18 cs.RO cs.SY eess.SP eess.SY 78%

RoboMNIST: A Multimodal Dataset for Multi-Robot Activity Recognition Using WiFi Sensing, Video, and Audio

Kian Behzad, Rojin Zandi, Elaheh Motamedi, Hojjat Salehinejad, Milad Siami

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18940 2025-02-17 cs.HC 78%

Amuse: Human-AI Collaborative Songwriting with Multimodal Inspirations

Yewon Kim, Sung-Ju Lee, Chris Donahue

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Published as a conference paper at CHI 2025. Project page: https://yewon-kim.com/amuse

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03862 2025-02-10 cs.HC 78%

Enhancing Deliberativeness: Evaluating the Impact of Multimodal Reflection Nudges

ShunYi Yeo, Zhuoqun Jiang, Anthony Tang, Simon Tangi Perrault

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02830 2025-02-06 cs.HC cs.LG q-bio.NC 78%

Multimodal Brain-Computer Interfaces: AI-powered Decoding Methodologies

Siyang Li, Hongbin Wang, Xiaoqing Chen, Dongrui Wu

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01801 2025-02-05 cs.HC 78%

MemPal: Leveraging Multimodal AI and LLMs for Voice-Activated Object Retrieval in Homes of Older Adults

Natasha Maniar, Samantha W. T. Chan, Wazeer Zulfikar, Scott Ren, Christine Xu, Pattie Maes

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15711 2025-02-05 cs.RO 78%

Robustifying Long-term Human-Robot Collaboration through a Multimodal and Hierarchical Framework

Peiqi Yu, Abulikemu Abuduweili, Ruixuan Liu, Changliu Liu

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12088 2025-02-05 cs.CY 78%

Mental-Perceiver: Audio-Textual Multi-Modal Learning for Estimating Mental Disorders

Jinghui Qin, Changsong Liu, Tianchi Tang, Dahuang Liu, Minghao Wang, Qianying Huang, Rumin Zhang

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments Accepted to AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05322 2025-01-14 cs.HC 78%

"What's Happening"- A Human-centered Multimodal Interpreter Explaining the Actions of Autonomous Vehicles

Xuewen Luo, Fan Ding, Ruiqi Chen, Rishikesh Panda, Junnyong Loo, Shuyun Zhang

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This paper has been accepted for presentation at WACV Workshop HAVI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12273 2024-12-31 cs.RO 78%

Multimodal Human-Autonomous Agents Interaction Using Pre-Trained Language and Visual Foundation Models

Linus Nwankwo, Elmar Rueckert

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00185 2024-12-03 astro-ph.IM astro-ph.GA 78%

Interactive Multimodal Integral Field Spectroscopy

Adrián García Riber, Rubén García-Benito, Francisco Serradilla

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 12 pages, 12 figures, 2 tables. Accepted for publication in RAS Techniques & Instruments (RASTI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18587 2024-11-28 cs.HC eess.SP q-bio.NC 78%

EEG-Based Analysis of Brain Responses in Multi-Modal Human-Robot Interaction: Modulating Engagement

Suzanne Oliver, Tomoko Kitago, Adam Buchwald, S. Farokh Atashzar

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments 9 pages, 7 figures. Submitted to IEEE TNSRE

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15590 2024-11-26 cs.LG cs.HC stat.ME 78%

From Complexity to Parsimony: Integrating Latent Class Analysis to Uncover Multimodal Learning Patterns in Collaborative Learning

Lixiang Yan, Dragan Gašević, Linxuan Zhao, Vanessa Echeverria, Yueqiao Jin, Roberto Martinez-Maldonado

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14147 2024-11-22 eess.IV 78%

Spiking neural networks: Towards bio-inspired multimodal perception in robotics

Katerina Maria Oikonomou, Vasiliki Balaska, Konstantinos A. Tsintotas, Christos N. Mavridis, Ioannis Kansizoglou, Antonios Gasteratos

专题命中 音频语音多模态 :multimodal(title);audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05342 2024-11-11 cs.RO 78%

Development of a Human-Robot Interaction Platform for Dual-Arm Robots Based on ROS and Multimodal Artificial Intelligence

Thanh Nguyen Canh, Ba Phuong Nguyen, Hong Quan Tran, Xiem HoangVan

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments In The 25th National Conference on Electronics, Communications and Information Technology (REV-ECIT 2022), Hanoi, Vietnam. in Vietnamese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19464 2024-11-05 cs.RO cs.AI cs.CV cs.SD eess.AS 78%

ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data

Zeyi Liu, Cheng Chi, Eric Cousineau, Naveen Kuppuswamy, Benjamin Burchfiel, Shuran Song

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.AI、eess.AS

Comments Conference on Robot Learning (CoRL) 2024; Project website: https://maniwav.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04279 2024-10-28 cs.SI 78%

MSEVA : A System for Multimodal Short Videos Emotion Visual Analysis

Qinglan Wei, Yaqi Zhou, Longhui Xiao, Yuan Zhang

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.01072 2024-10-28 cs.LG 78%

Interpretability for Multimodal Emotion Recognition using Concept Activation Vectors

Ashish Ramayee Asokan, Nidarshan Kumar, Anirudh Venkata Ragam, Shylaja S Sharath

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08578 2024-10-22 cs.HC eess.SP 78%

Dynamics of Collective Group Affect: Group-level Annotations and the Multimodal Modeling of Convergence and Divergence

Navin Raj Prabhu, Maria Tsfasman, Catharine Oertel, Timo Gerkmann, Nale Lehmann-Willenbrock

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏