arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4562 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4562 篇

2405.07930 2024-10-15 cs.MM cs.CV cs.LG cs.SD eess.AS 78%

Improving Multimodal Learning with Multi-Loss Gradient Modulation

Konstantinos Kontras, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00010 2024-10-02 eess.SP cs.LG 78%

PHemoNet: A Multimodal Network for Physiological Signals

Eleonora Lopez, Aurelio Uncini, Danilo Comminiello

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments The paper has been accepted at RTSI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06793 2024-09-25 cs.CR cs.IR cs.LG 78%

Adversarial Attacks to Multi-Modal Models

Zhihao Dou, Xin Hu, Haibo Yang, Zhuqing Liu, Minghong Fang

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments To appear in the ACM Workshop on Large AI Systems and Models with Privacy and Safety Analysis 2024 (LAMPS '24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15033 2024-09-24 cs.HC 78%

Immersed in my Ideas: Using Virtual Reality and Multimodal Interactions to Visualize Users' Ideas and Thoughts

Yunhao Xing, Jerrick Ban, Timothy D. Hubbard, Michael Villano, Diego Gomez-Zara

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 24 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07313 2024-08-15 cs.HC 78%

Exploring Large-Scale Language Models to Evaluate EEG-Based Multimodal Data for Mental Health

Yongquan Hu, Shuning Zhang, Ting Dang, Hong Jia, Flora D. Salim, Wen Hu, Aaron J. Quigley

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 6 pages; UbiComp Companion '24, Companion of the 2024 ACM International Joint Conference on Pervasive and Ubiquitous Computing, October 5--9, 2024}{Melbourne, VIC, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14116 2024-08-15 cs.RO cs.HC cs.LG 78%

Learning Multimodal Confidence for Intention Recognition in Human-Robot Interaction

Xiyuan Zhao, Huijun Li, Tianyuan Miao, Xianyi Zhu, Zhikai Wei, Aiguo Song

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15507 2024-06-21 cs.HC 78%

"May I Speak?": Multi-modal Attention Guidance in Social VR Group Conversations

Geonsun Lee, Dae Yeol Lee, Guan-Ming Su, Dinesh Manocha

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00841 2024-06-11 cs.RO 78%

Advantages of Multimodal versus Verbal-Only Robot-to-Human Communication with an Anthropomorphic Robotic Mock Driver

Tim Schreiter, Lucas Morillo-Mendez, Ravi T. Chadalavada, Andrey Rudenko, Erik Billing, Martin Magnusson, Kai O. Arras, Achim J. Lilienthal

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This paper has been accepted to the 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), which will be held in Busan, South Korea on August 28-31, 2023. For more information, please visit: https://ro-man2023.org/main

Journal ref 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19941 2024-05-31 cs.HC cs.CY 78%

Synthetic Patients: Simulating Difficult Conversations with Multimodal Generative AI for Medical Education

Simon N. Chu, Alex J. Goodell

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12353 2024-05-22 cs.LG 78%

TinyM$^2$Net-V3: Memory-Aware Compressed Multimodal Deep Neural Networks for Sustainable Edge Deployment

Hasib-Al Rashid, Tinoosh Mohsenin

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted at AAAI 2024 Workshop SAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15174 2024-04-12 cs.RO cs.HC 78%

LaMI: Large Language Models for Multi-Modal Human-Robot Interaction

Chao Wang, Stephan Hasler, Daniel Tanneberg, Felix Ocker, Frank Joublin, Antonello Ceravola, Joerg Deigmoeller, Michael Gienger

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03498 2024-04-05 cs.RO cs.HC 78%

Integrating Large Language Models with Multimodal Virtual Reality Interfaces to Support Collaborative Human-Robot Construction Work

Somin Park, Carol C. Menassa, Vineet R. Kamat

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 39 pages, 16 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19841 2024-04-01 cs.IR 78%

Dealing with Missing Modalities in Multimodal Recommendation: a Feature Propagation-based Approach

Daniele Malitesta, Emanuele Rossi, Claudio Pomo, Fragkiskos D. Malliaros, Tommaso Di Noia

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19095 2024-03-29 cs.CY 78%

Purposeful remixing with generative AI: Constructing designer voice in multimodal composing

Xiao Tan, Wei Xu, Chaoran Wang

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12609 2024-03-20 cs.LG 78%

SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition

Denis Dresvyanskiy, Maxim Markitantov, Jiawei Yu, Peitong Li, Heysem Kaya, Alexey Karpov

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract)

Comments 9 pages,

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12267 2024-03-14 cs.RO cs.HC cs.LG 78%

Continuous ErrP detections during multimodal human-robot interaction

Su Kyoung Kim, Michael Maurus, Mathias Trampler, Marc Tabie, Elsa Andrea Kirchner

专题命中 音频语音多模态 :multimodal(title,abstract)

Journal ref Int. Conf. Human-Computer Interaction (2023) 92-101

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07257 2024-03-04 cs.RO 78%

The Audio-Visual BatVision Dataset for Research on Sight and Sound

Amandine Brunetto, Sascha Hornauer, Stella X. Yu, Fabien Moutarde

专题命中 音频语音多模态 :audio-visual(title,abstract)

Comments Project page https://amandinebtto.github.io/Batvision-Dataset/ This version contains camera ready paper

Journal ref 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06306 2024-02-12 cs.IT eess.SP math.IT 78%

Multi-Modal Concurrent Transmission

Majid Nasiri Khormuji, Alberto Giuseppe Perotti, Qin Yi, Branislav Popovic

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments 6 pages, 4 figures, 1 table

Journal ref 2024 IEEE Wireless Communications and Networking Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10179 2023-12-19 cs.LG 78%

3FM: Multi-modal Meta-learning for Federated Tasks

Minh Tran, Roochi Shah, Zejun Gong

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09740 2023-12-18 cs.RO 78%

VITA: A Multi-modal LLM-based System for Longitudinal, Autonomous, and Adaptive Robotic Mental Well-being Coaching

Micol Spitale, Minja Axelsson, Hatice Gunes

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10170 2023-11-20 cs.LG 78%

Improving Unimodal Inference with Multimodal Transformers

Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04160 2023-11-08 cs.HC 78%

"Tell me about that church": Exploring the Design and User Experience of In-Vehicle Multi-modal Intuitive Interface in the Context of Driving Scenario

Yueteng Yu, Yan Zhang, Gary Burnett

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14118 2023-11-07 cs.LG 78%

MultiModN- Multimodal, Multi-Task, Interpretable Modular Networks

Vinitra Swamy, Malika Satayeva, Jibril Frej, Thierry Bossy, Thijs Vogels, Martin Jaggi, Tanja Käser, Mary-Anne Hartley

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted as a full paper at NeurIPS 2023 in New Orleans, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15247 2023-10-25 cs.SD cs.CV cs.LG cs.MM eess.AS 78%

SyncFusion: Multimodal Onset-synchronized Video-to-Audio Foley Synthesis

Marco Comunità, Riccardo F. Gramaccioni, Emilian Postolache, Emanuele Rodolà, Danilo Comminiello, Joshua D. Reiss

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01733 2023-10-04 eess.SY cs.SY 78%

Health Guardian: Using Multi-modal Data to Understand Individual Health

Vince S. Siu, Kuan Yu Hsieh, Italo Buleje, Takashi Itoh, Tian Hao, Ben Civjan, Nigel Hinds, Bing Dang, Jeffrey L. Rogers, Bo Wen

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments 10 pages, 6 figures

Journal ref IEEE International Conference on Digital Health (ICDH), 2023, pp. 65-74

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12249 2023-06-22 cs.SD cs.AI cs.IR cs.LG cs.MM eess.AS 78%

Knowledge-based Multimodal Music Similarity

Andrea Poltronieri

专题命中 音频语音多模态 :multimodal(title);分类 cs.AI、cs.MM、eess.AS

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03699 2023-05-24 cs.HC 78%

Multimodal User Authentication in Smart Environments: Survey of User Attitudes

Aishat Aloba, Sarah Morrison-Smith, Aaliyah Richlen, Kimberly Suarez, Yu-Peng Chen, Shaghayegh Esmaeili, Damon L. Woodard, Jaime Ruiz, Lisa Anthony

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 23 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12063 2023-05-23 cs.LG cs.HC 78%

Efficient Multimodal Neural Networks for Trigger-less Voice Assistants

Sai Srujana Buddi, Utkarsh Oggy Sarawgi, Tashweena Heeramun, Karan Sawnhey, Ed Yanosik, Saravana Rathinam, Saurabh Adya

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10588 2023-04-24 cs.RO cs.HC 78%

Detecting Worker Attention Lapses in Human-Robot Interaction: An Eye Tracking and Multimodal Sensing Study

Zhuangzhuang Dai, Jinha Park, Aleksandra Kaszowska, Chen Li

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04385 2023-04-12 cs.LG 78%

On Robustness in Multimodal Learning

Brandon McKinzie, Joseph Cheng, Vaishaal Shankar, Yinfei Yang, Jonathon Shlens, Alexander Toshev

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏