arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4559 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4559 篇

2308.03690 2024-09-12 cs.RO cs.AI 79%

Safe Multimodal Communication in Human-Robot Collaboration

Davide Ferrari, Andrea Pupa, Alberto Signoretti, Cristian Secchi

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref Human-Friendly Robotics 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02566 2024-09-05 cs.CV 79%

How Do You Perceive My Face? Recognizing Facial Expressions in Multi-Modal Context by Modeling Mental Representations

Florian Blume, Runfeng Qu, Pia Bideau, Martin Maier, Rasha Abdel Rahman, Olaf Hellwich

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments GCPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00012 2024-09-04 cs.HC cs.AI 79%

AVIN-Chat: An Audio-Visual Interactive Chatbot System with Emotional State Tuning

Chanhyuk Park, Jungbin Cho, Junwan Kim, Seongmin Lee, Jungsu Kim, Sanghoon Lee

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI

Journal ref Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI) Demo Track. 8763-8766 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17430 2024-08-29 eess.AS 79%

A Comprehensive Review and Taxonomy of Audio-Visual Synchronization Techniques for Realistic Speech Animation

Jose Geraldo Fernandes, Sinval Nascimento, Daniel Dominguete, André Oliveira, Lucas Rotsen, Gabriel Souza, David Brochero, Luiz Facury, Mateus Vilela, Hebert Costa, Frederico Coelho, Antônio P. Braga

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01760 2024-08-28 cs.CY cs.AI 79%

Trust and ethical considerations in a multi-modal, explainable AI-driven chatbot tutoring system: The case of collaboratively solving Rubik's Cube

Kausik Lakkaraju, Vedant Khandelwal, Biplav Srivastava, Forest Agostinelli, Hengtao Tang, Prathamjeet Singh, Dezhi Wu, Matt Irvin, Ashish Kundu

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted at 'Neural Conversational AI Workshop - What's left to TEACH (Trustworthy, Enhanced, Adaptable, Capable, and Human-centric) chatbots?' at ICML 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21136 2024-08-27 cs.CV 79%

MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls

Yuxuan Bian, Ailing Zeng, Xuan Ju, Xian Liu, Zhaoyang Zhang, Wei Liu, Qiang Xu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12558 2024-08-23 cs.MM 79%

Exploring the Role of Audio in Multimodal Misinformation Detection

Moyang Liu, Yukun Liu, Ruibo Fu, Zhengqi Wen, Jianhua Tao, Xuefei Liu, Guanjun Li

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08068 2024-08-19 cs.CL 79%

Large Language Models Meet Text-Centric Multimodal Sentiment Analysis: A Survey

Hao Yang, Yanyan Zhao, Yang Wu, Shilong Wang, Tian Zheng, Hongbo Zhang, Zongyang Ma, Wanxiang Che, Bing Qin

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments arXiv admin note: text overlap with arXiv:2210.14556 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06858 2024-08-14 eess.AS 79%

SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis

Osamu Take, Shinnosuke Takamichi, Kentaro Seki, Yoshiaki Bando, Hiroshi Saruwatari

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments 5 pages, accepted for INTERSPEECH 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04535 2024-08-13 eess.IV cs.AI 79%

Synchronous Multi-modal Semantic Communication System with Packet-level Coding

Yun Tian, Jingkai Ying, Zhijin Qin, Ye Jin, Xiaoming Tao

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05364 2024-08-13 cs.CV 79%

Spherical World-Locking for Audio-Visual Localization in Egocentric Videos

Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla, Anurag Kumar, Jacob Donley, Chao Li, Gunhee Kim, Vamsi Krishna Ithapu, Calvin Murdock

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07622 2024-08-12 cs.MM 79%

EMID: An Emotional Aligned Dataset in Audio-Visual Modality

Jialing Zou, Jiahao Mei, Guangze Ye, Tianyu Huai, Qiwei Shen, Daoguo Dong

专题命中 音频语音多模态 :audio-visual(title);cross-modal(abstract);分类 cs.MM

Comments Accepted by ACM MM workshop McGE 2023: Proceedings of the 1st International Workshop on Multimedia Content Generation and Evaluation: New Methods and Practice

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03185 2024-08-07 cs.CR cs.MM 79%

MaskAnyone Toolkit: Offering Strategies for Minimizing Privacy Risks and Maximizing Utility in Audio-Visual Data Archiving

Babajide Alamu Owoyele, Martin Schilling, Rohan Sawahn, Niklas Kaemer, Pavel Zherebenkov, Bhuvanesh Verma, Wim Pouw, Gerard de Melo

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12629 2024-08-07 cs.MM cs.CY cs.SI 79%

Television Discourse Decoded: Comprehensive Multimodal Analytics at Scale

Anmol Agarwal, Pratyush Priyadarshi, Shiven Sinha, Shrey Gupta, Hitkul Jangra, Ponnurangam Kumaraguru, Kiran Garimella

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments KDD 2024 [Updates for Camera Ready version]

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01708 2024-08-06 cs.CV 79%

AVESFormer: Efficient Transformer Design for Real-Time Audio-Visual Segmentation

Zili Wang, Qi Yang, Linsu Shi, Jiazhong Yu, Qinghua Liang, Fei Li, Shiming Xiang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10598 2024-08-06 eess.AS cs.SD 79%

Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge

Federico Costa, Miquel India, Javier Hernando

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Odyssey 2024: The Speaker and Language Recognition Workshop

Journal ref Proc. The Speaker and Language Recognition Workshop (Odyssey 2024), 266-273

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15308 2024-07-30 cs.CV 79%

AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset

Zhixi Cai, Shreya Ghosh, Aman Pankaj Adatia, Munawar Hayat, Abhinav Dhall, Tom Gedeon, Kalin Stefanov

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13660 2024-07-19 cs.LG cs.SD eess.AS 79%

CogniVoice: Multimodal and Multilingual Fusion Networks for Mild Cognitive Impairment Assessment from Spontaneous Speech

Jiali Cheng, Mohamed Elgaar, Nidhi Vakil, Hadi Amiri

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments INTERSPEECH 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06005 2024-07-09 cs.CV 79%

Advancing Automated Deception Detection: A Multimodal Approach to Feature Extraction and Analysis

Mohamed Bahaa, Mena Hany, Ehab E. Zakaria

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12853 2024-07-02 cs.CV 79%

Inconsistency-Aware Cross-Attention for Audio-Visual Fusion in Dimensional Emotion Recognition

G Rajasekhar, Jahangir Alam

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2403.19554

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15704 2024-06-25 cs.CV 79%

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, Yuxuan Wang, Chao Zhang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted at ICML 2024. arXiv admin note: substantial text overlap with arXiv:2310.05863

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15177 2024-06-24 cs.MM 79%

EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot

Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, Erik Cambria

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments ACL 2024 Demonstration Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14576 2024-06-24 eess.AS 79%

Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis

Kubilay Can Demir, Belen Lojo Rodriguez, Tobias Weise, Andreas Maier, Seung Hee Yang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 5 Pages, Interspeech 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12252 2024-06-19 cs.CL 79%

Language and Multimodal Models in Sports: A Survey of Datasets and Applications

Haotian Xia, Zhengbang Yang, Yun Zhao, Yuqing Wang, Jingxi Li, Rhys Tracy, Zhuangdi Zhu, Yuan-fang Wang, Hanjie Chen, Weining Shen

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10152 2024-06-17 cs.SD eess.AS 79%

Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition

Guinan Li, Jiajun Deng, Youjun Chen, Mengzhe Geng, Shujie Hu, Zhe Li, Zengrui Jin, Tianzi Wang, Xurong Xie, Helen Meng, Xunying Liu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted by Interspeech 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09556 2024-06-13 eess.SP cs.AI cs.IT math.IT 79%

Co-learning-aided Multi-modal-deep-learning Framework of Passive DOA Estimators for a Heterogeneous Hybrid Massive MIMO Receiver

Jiatong Bai, Feng Shu, Qinghe Zheng, Bo Xu, Baihua Shi, Yiwen Chen, Weibin Zhang, Xianpeng Wang

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06865 2024-06-12 cs.AI 79%

Eyeballing Combinatorial Problems: A Case Study of Using Multimodal Large Language Models to Solve Traveling Salesman Problems

Mohammed Elhenawy, Ahmed Abdelhay, Taqwa I. Alhadidi, Huthaifa I Ashqar, Shadi Jaradat, Ahmed Jaber, Sebastien Glaser, Andry Rakotonirainy

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06287 2024-06-10 cs.CV 79%

Hierarchical Augmentation and Distillation for Class Incremental Audio-Visual Video Recognition

Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07160 2024-06-04 cs.SD cs.LG eess.AS 79%

LLark: A Multimodal Instruction-Following Language Model for Music

Josh Gardner, Simon Durand, Daniel Stoller, Rachel M. Bittner

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments ICML camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16701 2024-05-28 cs.CV 79%

Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition

Tong Shi, Xuri Ge, Joemon M. Jose, Nicolas Pugeault, Paul Henderson

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Submitted to 27th International Conference of Pattern Recognition (ICPR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏