arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4549 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4549 篇

2305.03369 2023-05-08 cs.LG cs.AI cs.CL cs.MM 85%

The MuSe 2023 Multimodal Sentiment Analysis Challenge: Mimicked Emotions, Cross-Cultural Humour, and Personalisation

Lukas Christ, Shahin Amiriparian, Alice Baird, Alexander Kathan, Niklas Müller, Steffen Klug, Chris Gagne, Panagiotis Tzirakis, Eva-Maria Meßner, Andreas König, Alan Cowen, Erik Cambria, Björn W. Schuller

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CL、cs.AI、cs.MM

Comments Baseline paper for the 4th Multimodal Sentiment Analysis Challenge (MuSe) 2023, a workshop at ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10720 2026-08-12 cs.AI cs.CL cs.CV 新提交 85%

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Ex-Omni-2D:具备原生视觉存在的高表达全模态对话模型

Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo

专题命中 音频语音多模态 :omni-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 Ex-Omni-2D是一种全模态对话框架,可生成含文本、个性化语音与参考条件视频的协同响应,通过特定机制实现高效增量生成,在指定分辨率下达成1.293的端到端RTF,提供实用的质量效率平衡点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20720 2026-08-05 cs.HC 版本更新 85%

Beyond Text: Probing K-12 Educators' Perspectives and Ideas for Learning Opportunities Leveraging Multimodal Large Language Models

超越文本:探索K-12教育工作者对利用多模态大语言模型学习机会的视角和想法

Tiffany Tseng, Katelyn Lam, Tiffany Lin Fu, Alekhya Maram

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn)

AI总结 研究通过工作坊探讨K-12教育工作者对多模态大语言模型在教育中的应用看法,分析其面临的挑战与需求,提出两种用户导向的实施方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24786 2026-07-29 cs.IR cs.AI cs.MM cs.SD eess.AS 新提交 85%

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models

在大型视听检索模型中解锁空间定位

Hugo Malard, Michel Olvera, Sanjeel Parekh, Gaël Richard, Slim Essid, Stéphane Lathuilière

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.AI、cs.MM、eess.AS

AI总结 研究针对视听声源定位任务,利用大规模视听检索模型的潜在表示,引入LAIP框架,通过音频信息池化恢复局部空间信息,在相关数据集上取得领先性能,证明可从现有检索表示解锁定位,为检索和定位任务提供统一路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10299 2026-07-14 cs.LG 新提交 85%

Empowering Long-form Omni-modal Understanding with Robust Audio Perception

通过强大的音频感知增强长格式全模态理解

Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie

机构 * SAI, Shanghai Jiao Tong University(上海交通大学 上海人工智能研究院) Zhejiang University(浙江大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 音频语音多模态 :omni-modal(title,abstract);multimodal(abstract);audio-visual(abstract)

AI总结 为解决全模态理解不足问题,提出AVDC数据集及AVDC-QA-CoT数据集,利用现成模型标注视频,采用两阶段训练范式,在多下游任务实验中取得显著性能提升,推动全模态感知发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.06995 2026-06-11 cs.SD cs.CV cs.MM eess.AS 版本更新 85%

Benchmarking Cross-Domain Audio-Visual Deception Detection

跨域音视频欺骗检测基准测试

Xiaobao Guo, Zitong Yu, Nithish Muthuchamy Selvaraj, Bingquan Shen, Adams Wai-Kin Kong, Alex C. Kot

机构 * Rapid-Rich Object Search (ROSE) Lab and the College of Computing and Data Science, Nanyang Technological University (NTU)(快速丰富对象搜索(ROSE)实验室和南洋理工大学计算与数据科学学院) School of Computing and Information Technology and Dongguan Key Laboratory for Intelligence and Information Technology, Great Bay University(计算与信息科技学院和东莞智能与信息技术重点实验室,大湾大学) DSO National Laboratories(国防科学实验室) College of Computing and Data Science, Nanyang Technological University (NTU)(计算与数据科学学院,南洋理工大学) SMBU, Shenzhen 518172, China(深圳SMBU,越南河内VinUniversity,和新加坡NTU) VinUniversity, Hanoi 100000, Vietnam and NTU, Singapore

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

AI总结 提出首个跨域音视频欺骗检测基准,评估不同场景下的泛化能力,并设计MM-IDGM算法和Attention-Mixer融合方法提升性能。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10147 2026-06-10 cs.AI cs.CL cs.CV cs.SD 新提交 85%

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs

从感知到决策:多模态大语言模型中听觉与视觉感知的信息流

Wish Suharitdamrong, Muhammad Awais, Xiatian Zhu, Sara Atito

机构 * Surrey Institute for People-Centred AI (PAI)(萨里人本人工智能研究所) University of Surrey(萨里大学) Centre for Vision, Speech and Signal Processing (CVSSP)(视觉、语音和信号处理中心)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究多模态大语言模型(AVLLMs)中音频和视觉信息流的路径与整合机制,发现顺序流与并行流两种路由模式,并证明信息传递后可丢弃无关token以提升效率。

Comments 40 pages, 29 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30965 2026-06-01 eess.AS cs.AI cs.CL 85%

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

ImmersiveTTS:基于多模态扩散Transformer和领域特定表示对齐的环境感知文本转语音

Jun-Hak Yun, Seung-Bin Kim, Seong-Whan Lee

机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 提出ImmersiveTTS模型,通过多模态扩散Transformer和领域特定表示对齐,实现与环境音频自然融合的文本到语音生成。

Comments Accepted to ACL 2026 main conference. Code is available at https://github.com/jjunak-yun/ImmersiveTTS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11367 2026-05-26 cs.DC 85%

Efficient Distributed MLLM Training with Cornstarch

使用Cornstarch高效分布式多模态大语言模型训练

Insu Jang, Runyu Lu, Nikhil Bansal, Ang Chen, Mosharaf Chowdhury

专题命中 音频语音多模态 :MLLM(title,abstract);multimodal(abstract)

AI总结 提出Cornstarch框架,通过冻结感知流水线并行和令牌工作负载平衡的上下文并行,解决多模态大语言模型训练中的异构性问题,平均吞吐量提升2.26倍。

Comments ICML'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08762 2026-05-12 cs.SD cs.LG 85%

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

Omni-DeepSearch:一种基于音频的多模态深度搜索基准

Tao Yu, yiming ding, Shenghua Chai, Minghui Zhang, Zhongtian Luo, Xinming Wang, Xinlong Chen, Zhaolu Kang, Junhao Gong, Yuxuan Zhou, Haopeng Jin, Zhiqing Cui, Jiabing Yang, YiFan Zhang, Hongzhu Yi, Zheqi He, Xi Yang, Yan Huang, Liang Wang

机构 * CASIA UCAS BAAI Peking University(北京大学) Tsinghua University(清华大学)

专题命中 音频语音多模态 :omni-modal(title,abstract);multimodal(abstract);cross-modal(abstract)

AI总结 本文提出Omni-DeepSearch基准,用于评估基于音频的多模态深度搜索能力,通过多跳推理生成客观答案,结果显示该任务极具挑战性。

Comments 43 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10708 2026-04-28 cs.SD cs.AI cs.CV cs.MM 85%

Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing

Audio-Omni: 扩展多模态理解到多功能音频生成与编辑

Zeyue Tian, Binxin Yang, Zhaoyang Liu, Jiexuan Zhang, Ruibin Yuan, Hubery Yin, Qifeng Chen, Chen Li, Jing Lyu, Wei Xue, Yike Guo

机构 * Hong Kong University of Science and Technology(香港理工大学) WeChat Vision, Tencent Inc(微信视觉,腾讯公司) Peking University(北京大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 Audio-Omni提出首个端到端框架,统一音频生成与编辑,结合多模态理解能力,通过冻结的多模态大语言模型与可训练的扩散变换器实现高保真合成,并构建大规模数据集提升音频编辑性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12863 2026-04-08 cs.SD cs.AI cs.CV eess.AS 85%

Unified Cross-modal Translation of Score Images, Symbolic Music, and Performance Audio

统一的跨模态评分图像、符号音乐和表演音频翻译

Jongmin Jung, Dongmin Kim, Sihun Lee, Seola Cho, Hyungjoon Soh, Irmak Bukey, Chris Donahue, Dasaem Jeong

机构 * Department of Artificial Intelligence, Sogang University(西江大学人工智能系) Sogang Future Lab, Sogang University(西江大学未来实验室) Department of Physics Education, Seoul National University(首尔大学物理教育系) Computer Science Department, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Department of Art & Technology, Sogang University(西江大学艺术与技术系)

专题命中 音频语音多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、eess.AS

AI总结 本文提出统一模型,通过大规模数据集和模态分词实现多模态翻译,提升光学音乐识别的符号错误率至13.67%并实现评分图像条件音频生成。

Comments Submitted to IEEE Transactions on Audio, Speech and Language Processing (TASLPRO)

Journal ref IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pp. 1876-1891, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04229 2026-04-07 cs.MM cs.AI cs.CV cs.SD 85%

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning

层次化语义相关性感知的掩码自编码器用于无监督音频-视觉表示学习

Donghuo Zeng, Hao Niu, Masato Taya

机构 * KDDI Research, Inc., Saitama, Japan(KDDI研究所,埼玉,日本)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出HSC-MAE框架,通过三级表示层次强制语义一致性,结合教师-学生架构和多任务学习,提升无监督音频-视觉表示学习效果。

Comments 6 pages, 2 tables, 4 figures. Accepted by IEEE ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17234 2026-03-10 cs.MM cs.AI cs.CV 85%

Taming Modality Entanglement in Continual Audio-Visual Segmentation

驯服持续音频-视觉分割中的模态纠缠

Yuyang Hong, Qi Yang, Tao Zhang, Zili Wang, Zhaojin Fu, Kun Ding, Bin Fan, Shiming Xiang

机构 * School of Artificial Intelligence, UCAS(人工智能学院,UCAS) MAIS, Institute of Automation(自动化研究所MAIS) School of Intelligent Science and Technology, University of Science and Technolog Beijing(智能科学与技术学院,北京理工大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出CAVS任务和CMR框架,通过解决多模态语义漂移和共现混淆问题,提升持续音频-视觉分割性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04128 2026-03-05 cs.CV cs.AI cs.MM 85%

Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

Crab$^{+}$: 一种可扩展且统一的音频视觉场景理解模型,具有显式合作

Dongnuan Cai, Henghui Du, Chang Zhou, Xi Chen, Dan Guo, Hongyuan Zhang, Xuelong Li, Di Hu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学耿丽人工智能学院) Institute of Artificial Intelligence of China Telecom (TeleAI)(中国电信人工智能研究院) AI Technology Center, Online Video Business Unit, Tencent PCG(腾讯PCG在线视频业务单元AI技术中心) Hefei University of Technology(合肥工业大学) The University of Hong Kong(香港大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 Crab$^{+}$通过显式合作解决音频视觉任务异质性问题,实现更广泛的任务覆盖和优于单任务模型的性能表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24195 2026-03-02 cs.AI cs.CL cs.CV cs.LG 85%

Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume

多模态大语言模型的不确定性量化:基于不一致性调整的语义体积

Gregory Kang Ruey Lau, Hieu Dao, Nicole Kan Hui Lin, Bryan Kian Hsiang Low

机构 * Department of Computer Science, National University of Singapore(新加坡国立大学计算机科学系)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 UMPIRE是一种无需训练的多模态大语言模型不确定性量化框架,通过内部特征有效捕捉语义多样性和响应不一致性,提升错误检测和不确定性校准性能。

Comments Earlier versions presented at ICLR 2025 QUESTION workshop and ICML 2025 R2-FM workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14259 2026-01-22 cs.CV cs.AI cs.HC cs.LG cs.SD eess.AS 85%

A Cloud-Based Cross-Modal Transformer for Emotion Recognition and Adaptive Human-Computer Interaction

基于云的跨模态Transformer用于情绪识别和自适应人机交互

Ziwen Zhong, Zhitao Shu, Yue Zhao

专题命中 音频语音多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、eess.AS

AI总结 本文提出基于云的跨模态Transformer框架,通过整合多模态信号提升情绪识别的鲁棒性和泛化能力,实现高效、实时的自适应人机交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01185 2025-12-19 cs.CR 85%

DefenSee: Dissecting Threat from Sight and Text -- A Multi-View Defensive Pipeline for Multi-modal Jailbreaks

DefenSee:从视觉和文本中解构威胁——一种多视图防御管道用于多模态对抗突破

Zihao Wang, Kar Wai Fok, Vrizlynn L. L. Thing

专题命中 音频语音多模态 :multi-modal(title,abstract);MLLM(abstract);cross-modal(abstract)

AI总结 DefenSee通过图像变体转录和跨模态一致性检查,提供一种多模态防御方法,有效降低多模态对抗攻击的成功率,提升模型鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19877 2025-12-12 cs.MM cs.CV cs.LG eess.AS 85%

It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models

它能听,也能看 too:通过将视觉理解整合到音频语言模型中构建多模态大语言模型以检测抑郁症

Xiangyu Zhao, Yaling Shen, Yiwen Jiang, Zimu Wang, Jiahe Liu, Maxmartwell H Cheng, Guilherme C Oliveira, Robert Desimone, Dominic Dwyer, Zongyuan Ge

机构 * Monash University(墨尔本大学) Massachusetts Institute of Technology(麻省理工学院) The University of Melbourne(墨尔本大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

AI总结 本文提出了一种多模态大语言模型框架,通过整合视觉理解到音频语言模型中,提升抑郁症检测的准确性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00120 2025-12-02 cs.SD cs.AI cs.CV cs.LG cs.MM 85%

Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment

Art2Music: 通过多模态情感对齐生成艺术图像的音乐

Jiaying Hong, Ting Zhu, Thanet Markchom, Huizhi Liang

机构 * School of Computing Newcastle University(计算学院新castle大学) Department of Computer Science University of Reading(计算机科学系阅读大学)

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 Art2Music通过多模态情感对齐生成艺术图像的音乐,利用轻量级跨模态框架提升音乐生成的感知自然性和频谱保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18698 2025-11-25 cs.SD cs.AI cs.CV cs.LG cs.MM 85%

Multimodal Real-Time Anomaly Detection and Industrial Applications

多模态实时异常检测与工业应用

Aman Verma, Keshav Samdani, Mohd. Samiuddin Shafi

机构 * Department of Electrical Engineering IIT Bombay Mumbai, India(电气工程系 印度理工学院Bombay 墨尔本)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出了一种多模态实时异常检测系统,结合视频和音频处理,通过多模型音频集合、混合物体检测和双向跨模态注意力机制,提升了准确性与工业适用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10016 2025-11-25 cs.MM cs.AI cs.CL 85%

A Survey of Generative Categories and Techniques in Multimodal Generative Models

多模态生成模型中生成类别的综述

Longzhen Han, Awes Mubarak, Almas Baimagambetov, Nikolaos Polatidis, Thar Baker

机构 * School of Architecture, Technology and Engineering, University of Brighton, Lewes Road, BN2 4GJ(建筑、科技与工程学院,布里顿大学,勒斯路,BN2 4GJ) University of Khorfakkan, Sharjah(柯法克坎大学,沙迦)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 本文综述多模态生成模型中生成类别的发展,分析了跨模态能力的基础技术,并探讨了评估框架与治理机制以提升系统通用性和安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17890 2025-11-25 cs.CV cs.AI cs.MM 85%

Decoupled Audio-Visual Dataset Distillation

解耦音频-视觉数据集蒸馏

Wenyuan Li, Guang Li, Keisuke Maeda, Takahiro Ogawa, Miki Haseyama

机构 * Hokkaido University(北海道大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 DAVDD通过解耦音频-视觉表示提升数据集蒸馏效果,利用预训练库和轻量解耦器实现稳定特征提取与模态隔离。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26825 2025-11-03 cs.SD cs.CV cs.MM eess.AS 85%

Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling

Jiarong Du, Zhan Jin, Peijun Yang, Juan Liu, Zhuo Li, Xin Liu, Ming Li

机构 * School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) School of Computer Science, Wuhan University(武汉大学计算机学院) Suzhou Municipal Key Laboratory of Multimodal Intelligent Systems, Digital Innovation Research Center, Duke Kunshan University(多模态智能系统苏州市级重点实验室、杜克昆山大学数字创新研究中心) Hardware Engineering System, OPPO(OPPO硬件工程系统)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11760 2025-10-15 cs.SD cs.AI cs.CV cs.MM 85%

Audio-Guided Visual Perception for Audio-Visual Navigation

Yi Wang, Yinfeng Yu, Fuchun Sun, Liejun Wang, Wendong Zheng

机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing(丝绸之路多语种认知计算国际联合实验室) Tsinghua University(清华大学) Tianjin University of Technology(天津工业大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Main paper (6 pages). Accepted for publication by International Conference on Virtual Reality and Visualization 2025 (ICVRV 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22744 2025-09-30 eess.AS cs.AI cs.MM cs.SD 85%

Index-MSR: A high-efficiency multimodal fusion framework for speech recognition

Jinming Chen, Lu Wang, Zheshu Song, Wei Deng

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM、eess.AS

Comments Submit to icassp 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13255 2025-09-24 cs.CL cs.AI cs.IR cs.LG cs.MM 85%

Automating Steering for Safe Multimodal Large Language Models

Lyucheng Wu, Mengru Wang, Ziwen Xu, Tri Cao, Nay Oo, Bryan Hooi, Shumin Deng

机构 * Zhejiang University(浙江大学) Zhejiang University - Ant Group Joint Lab of Knowledge Graph(浙江大学-蚂蚁集团知识图谱联合实验室) National University of Singapore, NUS-NCS Joint Lab(新加坡国立大学NUS-NCS联合实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments EMNLP 2025 Main Conference. 23 pages (8+ for main); 25 figures; 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15661 2025-09-22 cs.SD cs.AI cs.CL eess.AS 85%

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

Qiaolin Wang, Xilin Jiang, Linyang He, Junkai Wu, Nima Mesgarani

机构 * Columbia University(哥伦比亚大学) University of Washington(华盛顿大学)

专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01284 2025-09-15 cs.MM cs.CV cs.SD eess.AS eess.IV 85%

Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework

Liuyuan Wen

机构 * School of Physical Sciences University of Science and Technology of China(中国科学技术大学物理科学学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted to BMVC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12226 2025-09-09 cs.CL cs.AI cs.CV cs.LG 85%

AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu

机构 * Fudan University(复旦大学) Multimodal Art Projection Research Community(多模态艺术投影研究社区) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 28 pages, 16 figures, under review, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏