arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

共收录 495 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 音视频/视觉语言融合 495 篇

2603.01382 2026-03-03 cs.SD cs.CL 50%

End-to-End Simultaneous Dysarthric Speech Reconstruction with Frame-Level Adaptor and Multiple Wait-k Knowledge Distillation

端到端的同时口吃语音重建与帧级适配模块及多视图知识蒸馏

Minghui Wu, Haitao Tang, Jiahuan Fan, Ruizhi Liao, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) iFlytek Co., Ltd.(科大讯飞股份有限公司)

专题命中 音视频/视觉语言融合 :information fusion(abstract)

AI总结 本文提出端到端的同时口吃语音重建系统,通过帧级适配模块和多视图知识蒸馏模块,提升重建语音的鲁棒性和可懂度。

Comments Submitted to 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)

Journal ref 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Singapore, 2025, pp. 1092-1097

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22961 2026-03-03 eess.AS 50%

Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models

为统一多模态语音识别适应语音基础模型与大语言模型

Jing-Xuan Zhang, Genshun Wan, Jin Li, Jianqing Gao, Duo Zhao, Zhen-Hua Ling

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 本文提出UASR-LLM框架,通过大语言模型与语音基础模型结合,实现多模态语音识别的统一优化与性能提升。

Comments 10 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23300 2026-02-27 cs.CL eess.AS 50%

A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations

一种用于对话中多模态情绪识别的专家混合模型

Soumya Dutta, Smruthi Balaji, Sriram Ganapathy

机构 * LEAP Lab, Department of Electrical Engineering(LEAP实验室,电气工程系) Microsoft(微软)

专题命中 音视频/视觉语言融合 :information fusion(abstract)

AI总结 MiSTER-E通过专家混合框架提升对话中多模态情绪识别的准确率,实现跨模态一致性与融合。

Comments Accepted to Elsevier Computer Speech and Language. 30 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09294 2026-02-11 cs.CE 50%

BrainTAP: Brain Disorder Prediction with Adaptive Distill and Selective Prior Integration

BrainTAP: 基于自适应蒸馏和选择性先验整合的脑部疾病预测

Zhenyu Lei, Aiying Zhang, Song Wang, Han Fan, Jundong Li

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 BrainTAP通过自适应蒸馏和选择性先验整合,提升脑部疾病预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04920 2026-02-06 cs.LG cs.SD 50%

CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal Learning

CyIN:循环信息潜在空间用于连接完整与不完整多模态学习

Ronghao Lin, Qiaolin He, Sijie Mai, Ying Zeng, Aolin Xiong, Li Huang, Yap-Peng Tan, Haifeng Hu

机构 * School of Electronics and Information Technology, Sun Yat-Sen University(中山大学电子与信息学院) School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电气与电子工程学院) School of Computer Science, South China Normal University(华南师范大学计算机科学学院) Desay SV Automotive Co., Ltd(德赛股份有限公司) Pazhou Laboratory(琶洲实验室)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 CyIN通过构建循环信息潜在空间,解决多模态学习中完整与不完整数据之间的性能差距,实现统一优化。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00914 2026-02-03 cs.CL cs.AI cs.CY cs.SD eess.AS 50%

A Baseline Multimodal Approach to Emotion Recognition in Conversations

一种用于对话中情感识别的基线多模态方法

Víctor Yeste, Rodrigo Rivas-Arévalo

机构 * School of Science, Engineering and Design, Universidad Europea de Valencia(科学、工程与设计学院,欧洲大学 Valencia)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 本文提出了一种基于Transformer文本分类器和自监督语音模型的多模态基线方法,用于对话中情感识别,并通过实验展示了多模态融合的优势。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16124 2025-12-30 cs.HC cs.LG 50%

ZIA: A Theoretical Framework for Zero-Input AI

ZIA:零输入AI的理论框架

Aditi De

机构 * Indian Institute of Technology Roorkee(印度理工学院罗奥基分校)

专题命中 音视频/视觉语言融合 :multi-modal fusion(abstract)

AI总结 ZIA提出了一种基于多模态融合的零输入AI框架,通过整合生物信号和上下文数据实现前瞻性意图预测,提升实时推理效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17282 2025-12-04 cs.AI cs.SD 50%

ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection

ERF-BA-TFD+: 一种用于音频视觉深度伪造检测的多模态模型

Xin Zhang, Jiaming Chu, Jian Zhao, Yuchu Jiang, Xu Yang, Lei Jin, Chi Zhang, Xuelong Li

专题命中 音视频/视觉语言融合 :audio-visual fusion(abstract)

AI总结 ERF-BA-TFD+通过结合增强接收场和音频视觉融合,提出了一种多模态深度伪造检测模型,在DDL-AV数据集上实现了最先进的检测性能。

Comments The paper is withdrawn after discovering a flaw in the theoretical derivation presented in Section Method. The incorrect step leads to conclusions that are not supported by the corrected derivation. We plan to reconstruct the argument and will release an updated version once the issue is fully resolved

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19509 2025-11-26 cs.LG 50%

TouchFormer: A Robust Transformer-based Framework for Multimodal Material Perception

TouchFormer: 一种基于变换器的鲁棒多模态材料感知框架

Kailin Lyu, Long Xiao, Jianing Zeng, Junhao Dong, Xuexin Liu, Zhuojun Zou, Haoyue Yang, Lin Shu, Jie Hao

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 TouchFormer通过模态自适应门控和注意力机制提升多模态材料感知的鲁棒性,并在细粒度分类任务中实现性能提升。

Comments 9 pages, 7 figures, Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02113 2025-11-11 cs.IR 50%

Enhancing Multimodal Recommendations with Vision-Language Models and Information-Aware Fusion

Hai-Dang Kieu, Min Xu, Thanh Trung Huynh, Dung D. Le

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07841 2025-11-04 cs.NI cs.LG 50%

Task-Oriented Multimodal Token Transmission in Resource-Constrained Multiuser Networks

Junhe Zhang, Wanli Ni, Pengwei Wang, Dongyu Wang

专题命中 音视频/视觉语言融合 :information fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27091 2025-11-03 cs.LG cs.AI quant-ph 50%

QiNN-QJ: A Quantum-inspired Neural Network with Quantum Jump for Multimodal Sentiment Analysis

Yiwei Chen, Kehuan Yan, Yu Pan, Daoyi Dong

机构 * School of Engineering, Yunnan University(云南大学工程学院) College of Computer and Data Science, Fuzhou University(福州大学计算机与数据科学学院) Institute of Cyber-Systems and Control, College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院智能系统与控制研究所) Australian Artificial Intelligence Institute, Faculty of Engineering and Information Technology, University of Technology Sydney(新南威尔士大学人工智能研究所)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23640 2025-10-29 cs.LG cs.AI 50%

Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning

Zihao Jing, Yan Sun, Yan Yi Li, Sugitha Janarthanan, Alana Deng, Pingzhao Hu

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23273 2025-10-28 cs.LG cs.AI q-bio.QM 50%

A Novel Framework for Multi-Modal Protein Representation Learning

Runjie Zheng, Zhen Wang, Anjie Qiao, Jiancong Xie, Jiahua Rao, Yuedong Yang

机构 * School of Computer Science and Engineering, Sun Yat-sen University (SYSU)(计算机科学与工程学院,中山大学)

专题命中 音视频/视觉语言融合 :information fusion(abstract)

Comments 35 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17289 2025-10-21 cs.CL 50%

Addressing Antisocial Behavior in Multi-Party Dialogs Through Multimodal Representation Learning

Hajar Bakarou, Mohamed Sinane El Messoussi, Anaïs Ollagnier

机构 * Universit\'e C \ te d'Azur, CNRS, Inria, I3S Sophia Antipolis France Universit\'e C \ te d'Azur, CNRS, Inria, I3S

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08802 2025-10-13 cs.LG 50%

Edu-EmotionNet: Cross-Modality Attention Alignment with Temporal Feedback Loops

S M Rafiuddin

机构 * Department of Computer Science Oklahoma State University Stillwater, Oklahoma, USA(计算机科学系 奥克拉荷马州立大学 斯蒂尔沃特 奥克拉荷马州 美国)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments 6 Pages, 6 Figures, 3 Tables, Accepted as a Regular Research paper at ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07282 2025-09-26 eess.AS 50%

Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild

Jing-Tong Tzeng, Bo-Hao Su, Ya-Tse Wu, Hsing-Hang Chou, Chi-Chun Lee

专题命中 音视频/视觉语言融合 :multi-modal fusion(abstract)

Comments Proceedings of Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18706 2025-09-24 cs.HC 50%

M4SER: Multimodal, Multirepresentation, Multitask, and Multistrategy Learning for Speech Emotion Recognition

Jiajun He, Xiaohan Shi, Cheng-Hung Hu, Jinyi Mi, Xingfeng Li, Tomoki Toda

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments Accepted by IEEE Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15667 2025-09-22 cs.CL cs.SD eess.AS 50%

VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion

Dimitrios Damianos, Leon Voukoutis, Georgios Paraskevopoulos, Vassilis Katsouros

机构 * Institute for Speech and Language Processing, Athena Research Center, Greece(语音与语言处理研究所,亚特兰蒂斯研究中心,希腊)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10307 2025-09-16 cs.IR 50%

CROSSAN: Towards Efficient and Effective Adaptation of Multiple Multimodal Foundation Models for Sequential Recommendation

Junchen Fu, Yongxin Ni, Joemon M. Jose, Ioannis Arapakis, Kaiwen Zheng, Youhua Li, Xuri Ge

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13072 2025-08-19 cs.AI 50%

A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis

Yuting Zhang, Tiantian Geng, Luoying Hao, Xinxing Cheng, Alexander Thorley, Xiaoxia Wang, Wenqi Lu, Sandeep S Hothi, Lei Wei, Zhaowen Qiu, Dipak Kotecha, Jinming Duan

机构 * School of Computer Science, University of Birmingham, Birmingham, UK Department of Cardiovascular Sciences, University of Birmingham, Birmingham, UK NIHR Birmingham Biomedical Research Centre West Midlands NHS Secure Data Environment, University Hospitals Birmingham NHS Foundation Trust, Birmingham, UK Department of Computing Mathematics, Manchester Metropolitan University, Manchester, UK Department of Cardiology, Heart Lung Centre, Royal Wolverhampton NHS Trust, Wolverhampton, UK Department of Cardiovascular Surgery, The First Affiliated Hospital with Nanjing Medical University , Nanjing,China College of Computer Control Engineering, Northeast Forestry University, Harbin, China Julius Center, University Medical Center Utrecht, the Netherlands Data Sciences, University of Manchester, Manchester, UK

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02133 2025-08-19 cs.HC 50%

Hierarchical MoE: Continuous Multimodal Emotion Recognition with Incomplete and Asynchronous Inputs

Yitong Zhu, Lei Han, Guanxuan Jiang, PengYuan Zhou, Yuyang Wang

专题命中 音视频/视觉语言融合 :information fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09105 2025-08-14 cs.AI 50%

SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling

Shixuan Sun, Siyuan Liang, Ruoyu Chen, Jianjie Huang, Jingzhi Li, Xiaochun Cao

机构 * Sun Yat-Sen University(孙中山大学) Nanyang Technological University(南洋理工大学) University of Chinese Academy of Science(中国科学院大学) Zhongguancun Academy(中关村学院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20417 2025-07-29 cs.SD cs.CR eess.AS 50%

Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection

Yassine El Kheir, Arnab Das, Enes Erdem Erdogan, Fabian Ritter-Guttierez, Tim Polzehl, Sebastian Möller

机构 * Speech and Language Technology, DFKI, Germany(DFKI语音与语言技术) Quality and Usability Lab, Technical University of Berlin, Germany(柏林技术大学可用性实验室) AI Team, Gretchen AI, Germany(Gretchen AI人工智能团队) Nanyang Technological University, Singapore(南洋理工大学)

专题命中 音视频/视觉语言融合 :hybrid fusion(abstract)

Comments ACCEPTED WASPAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06247 2025-07-10 physics.flu-dyn physics.ins-det 50%

FED-PV: A Large-Scale Synthetic Frame/Event Dataset for Particle-Based Velocimetry

Fan Wu, Xiang Feng, Aoyu Zhang, Yong Lee

专题命中 音视频/视觉语言融合 :information fusion(abstract)

Comments This work has been accepted as a conference paper at the 16th International Symposium on Particle Image Velocimetry (ISPIV 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12151 2025-07-08 cs.LG cs.AI 50%

Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment Analysis

Miaosen Luo, Yuncheng Jiang, Sijie Mai

机构 * School of Computer Science, South China Normal University(华南师范大学计算机学院)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22446 2025-07-01 cs.LG cs.AI 50%

EAGLE: Efficient Alignment of Generalized Latent Embeddings for Multimodal Survival Prediction with Interpretable Attribution Analysis

Aakash Tripathi, Asim Waqas, Matthew B. Schabath, Yasin Yilmaz, Ghulam Rasool

机构 * Dept. of Machine Learning Moffitt Cancer Center(机器学习系莫菲特癌症中心) Dept. of Cancer Epidemiology Moffitt Cancer Center(癌症流行病学系莫菲特癌症中心) Dept. of Electrical Engineering University of South Florida(电气工程系佛罗里达州立大学)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07895 2025-06-23 cs.LG cs.AI 50%

Representation Learning with Mutual Influence of Modalities for Node Classification in Multi-Modal Heterogeneous Networks

Jiafan Li, Jiaqi Zhu, Liang Chang, Yilin Li, Miaomiao Li, Yang Wang, Hongan Wang

机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) Binzhou Institute of Technology(滨州职业技术学院)

专题命中 音视频/视觉语言融合 :multi-modal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10176 2025-05-16 cs.NE 50%

Incorporating brain-inspired mechanisms for multimodal learning in artificial intelligence

Xiang He, Dongcheng Zhao, Yang Li, Qingqun Kong, Xin Yang, Yi Zeng

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments The manuscript is under review and the code is available at https://github.com/Brain-Cog-Lab/IEMF

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00422 2025-05-02 cs.LG cs.CL 50%

Toward Automated Regulatory Decision-Making: Trustworthy Medical Device Risk Classification with Multimodal Transformers and Self-Training

Yu Han, Aaron Ceross, Jeroen H. M. Bergmann

机构 * Institute of Biomedical Engineering, Department of Engineering Science, University of Oxford(生物医学工程研究所,工程科学系,牛津大学) Birmingham Law School, University of Birmingham(伯明翰法学院,伯明翰大学)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏