arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2602.19832 2026-02-24 cs.CV 83%

M3S-Net: Multimodal Feature Fusion Network Based on Multi-scale Data for Ultra-short-term PV Power Forecasting

M3S-Net:基于多尺度数据的多模态特征融合网络用于超短期光伏功率预测

Penghui Niu, Taotao Cai, Suqi Zhang, Junhua Gu, Ping Zhang, Qiqi Liu, Jianxin Li

机构 * School of Artificial Intelligence, Hebei University of Technology(人工智能学院,河北工业大学) University of Southern Queensland(南方昆士兰大学) School of Information Engineering, Tianjin University of Commerce(信息工程学院,天津商业大学) Hebei Province Key Laboratory of Big Data Calculation, Hebei University of Technology(大数据计算河北重点实验室,河北工业大学) General AI Lab, School of Engineering, Westlake University(通用人工智能实验室,西湖大学) School of Business and Law, Edith Cowan University(商学院和法学院,埃迪斯科文大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 M3S-Net通过多尺度数据融合和跨模态Mamba交互模块,提升超短期光伏功率预测的精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19505 2026-02-24 cs.CV 83%

Test-Time Computing for Referring Multimodal Large Language Models

测试时计算用于指代多模态大语言模型

Mingrui Wu, Hao Chen, Jiayi Ji, Xiaoshuai Sun, Zhiyuan Liu, Liujuan Cao, Ming-Ming Cheng, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算重点实验室,中国教育部,厦门大学) VCIP, CS, Nankai University(VCIP,计算机科学,南开大学) Tsinghua University(清华大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 ControlMLLM++通过注入可学习的视觉提示实现多模态大语言模型的测试时适应,提升细粒度视觉推理能力。

Comments arXiv admin note: substantial text overlap with arXiv:2407.21534

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17196 2026-02-20 cs.CV 83%

EntropyPrune: Matrix Entropy Guided Visual Token Pruning for Multimodal Large Language Models

EntropyPrune: 基于矩阵熵的视觉令牌修剪用于多模态大语言模型

Yahong Wang, Juncheng Wu, Zhangkai Ni, Chengmei Yang, Yihang Liu, Longzhen Yang, Yuyin Zhou, Ying Wen, Lianghua He

机构 * School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) University of California, Santa Cruz(加州大学圣克ruz分校) East China Normal University(华东师范大学) Shanghai Eye Disease Prevention and Treatment Center(上海眼病预防与治疗中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 EntropyPrune通过矩阵熵指导的视觉令牌修剪方法,提升多模态大语言模型的推理效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08126 2026-02-13 cs.CV 83%

MambaFusion: Adaptive State-Space Fusion for Multimodal 3D Object Detection

MambaFusion:多模态3D目标检测的自适应状态空间融合

Venkatraman Narayanan, Bala Sai, Rahul Ahuja, Pratik Likhar, Varun Ravi Kumar, Senthil Yogamani

机构 * Qualcomm Technologies, Inc(高通技术公司) Qualcomm India Private Limited(高通印度私人有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 MambaFusion通过结合选择性状态空间模型与窗口化变换器,实现高效且可靠的多模态3D目标检测,提升自动驾驶系统的感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22463 2026-02-12 cs.MM 83%

Orthogonal Disentanglement with Projected Feature Alignment for Multimodal Emotion Recognition in Conversation

正交解缠与投影特征对齐用于对话中多模态情绪识别

Xinyi Che, Wenbo Wang, Jian Guan, Qijun Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

AI总结 本文提出OD-PFA框架,通过正交解缠与投影特征对齐技术提升对话中多模态情绪识别性能。

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05885 2026-02-09 cs.IR cs.AI 83%

An item is worth one token in Multimodal Large Language Models-based Sequential Recommendation

在基于多模态大语言模型的序列推荐中,一个项目相当于一个标记

Qiyong Zhong, Jiajie Su, Ming Yang, Yunshan Ma, Xiaolin Zheng, Chaochao Chen

机构 * Zhejiang University(浙江大学) Singapore Management University(新加坡国立管理学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 Speeder通过多模态表示压缩、模态感知渐进优化和序列位置感知增强,提升多模态大语言模型在序列推荐中的效率与效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05729 2026-02-06 cs.CV cs.LG 83%

Adaptive Global and Fine-Grained Perceptual Fusion for MLLM Embeddings Compatible with Hard Negative Amplification

自适应全局与细粒度感知融合用于兼容硬负样本放大的人脸嵌入

Lexiang Hu, Youze Xue, Dian Li, Gang Liu, Zhouchen Lin

机构 * State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(人工智能国家重点实验室,智能科学与技术学院,北京大学) Institute for Artificial Intelligence, Peking University(人工智能研究院,北京大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出AGFF-Embed方法,通过自适应融合全局和细粒度语义信息,提升多模态嵌入在一般和细粒度理解上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22936 2026-02-06 cs.CV 83%

PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models

PPE:用于多模态大语言模型中token压缩的位置保持嵌入

Mouxiao Huang, Borui Jiang, Dehua Zheng, Hailin Hu, Kai Han, Xinghao Chen

机构 * Huawei Technologies(华为技术有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 PPE通过保持位置信息提升多模态大语言模型的token压缩效率和性能

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15098 2026-02-04 cs.LG cs.AI 83%

How Intermodal Interaction Affects the Performance of Deep Multimodal Fusion for Mixed-Type Time Series

多模态交互如何影响深度多模态融合在混合类型时间序列中的性能

Simon Dietz, Thomas Altstidl, Dario Zanca, Björn Eskofier, An Nguyen

机构 * FAU Erlangen-Nürnberg(弗赖堡大学埃尔朗根-纽伦堡分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文研究了多模态交互对深度多模态融合在混合类型时间序列预测中的影响,通过三种融合类型和五种融合方法的比较,揭示了交互强度和方向对融合策略选择的关键作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02784 2026-02-04 cs.LG cs.AI 83%

Cross-Temporal Attention Fusion (CTAF) for Multimodal Physiological Signals in Self-Supervised Learning

跨时间注意力融合(CTAF)用于自监督学习中的多模态生理信号

Arian Khorasani, Théophile Demazure

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 CTAF通过时间感知的交叉注意力机制,实现多模态生理信号在自监督学习中的高效融合,提升匹配对的余弦边距和跨模态检索性能,同时保持高准确率并减少标签依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21280 2026-02-02 cs.CV 83%

Token Entropy Regularization for Multi-modal Antenna Affiliation Identification

基于令牌熵正则化的多模态天线归属识别

Dong Chen, Ruoyu Li, Xinyan Zhang, Jialei Xu, Ruosen Zhao, Zhikang Zhang, Lingyun Li, Zizhuang Wei

机构 * Huawei(华为) The University of Hong Kong(香港大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出基于令牌熵正则化的多模态天线归属识别方法,通过融合视频、几何特征和PCI信号,提升通信网络优化效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22020 2026-01-30 cs.LG cs.CV 83%

Visual-Guided Key-Token Regularization for Multimodal Large Language Model Unlearning

多模态大语言模型去敏中的视觉引导关键标记正则化

Chengyi Cai, Zesheng Ye, Peike Li, Bo Han, Jianzhong Qi, Feng Liu

机构 * The University of Melbourne(墨尔本大学) Google Research(谷歌研究) Hong Kong Baptist University(香港 Baptist 大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出视觉引导的关键标记正则化方法,用于多模态大语言模型的去敏,通过信息熵定义关键标记并利用梯度重新加权提升去敏效果,实验表明能有效减少遗忘并保持响应一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22009 2026-01-30 cond-mat.mtrl-sci cs.AI cs.LG physics.comp-ph 83%

MEIDNet: Multimodal generative AI framework for inverse materials design

MEIDNet: 多模态生成式AI框架用于反向材料设计

Anand Babu, Rogério Almeida Gouvêa, Pierre Vandergheynst, Gian-Marco Rignanese

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 MEIDNet通过多模态生成式AI框架实现反向材料设计,利用对比学习提升学习效率,生成低带隙钙钛矿结构并验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21547 2026-01-30 cs.LG cs.AI 83%

Multi-Modal Time Series Prediction via Mixture of Modulated Experts

多模态时间序列预测 via 专家混合

Lige Zhang, Ali Maatouk, Jialin Chen, Leandros Tassiulas, Rex Ying

机构 * Duke Kunshan University, China(杜克昆山大学) Yale University, USA(耶鲁大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出专家调制方法,通过条件控制专家行为实现多模态时间序列预测的高效跨模态控制。

Comments 26 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20258 2026-01-29 cs.CV cs.LG 83%

Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization

模态平衡的协同蒸馏用于多模态领域泛化

Xiaohan Wang, Zhangtao Cheng, Ting Zhong, Leiting Chen, Fan Zhou

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MBCD框架,通过自适应模态丢弃、梯度一致性约束和跨模态蒸馏,解决多模态领域泛化中模态不平衡问题,提升模型泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15139 2026-01-27 cs.CV 83%

Unified Cross-Modal Attention-Mixer Based Structural-Functional Connectomics Fusion for Neuropsychiatric Disorder Diagnosis

统一的跨模态注意力-混合器基于结构-功能连接组融合的神经精神疾病诊断

Badhan Mazumder, Lei Wu, Vince D. Calhoun, Dong Hye Ye

机构 * Department of Computer Science, Georgia State University(计算机科学系,佐治亚州立大学) Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, and Emory University(跨机构神经影像与数据科学转化研究中心(TReNDS),佐治亚州立大学、佐治亚理工学院和埃默里大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出ConneX方法,通过统一的跨模态注意力和MLP-Mixer实现结构-功能连接组的多模态融合,提升神经精神疾病诊断性能。

Comments Published in the Proceedings of the 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2025). IEEE Xplore. DOI: 10.1109/EMBC58623.2025.11254194

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16381 2026-01-26 cs.CV 83%

VTFusion: A Vision-Text Multimodal Fusion Network for Few-Shot Anomaly Detection

VTFusion: 一种面向少样本异常检测的视觉-文本多模态融合网络

Yuxin Jiang, Yunkang Cao, Yuqi Cheng, Yiheng Zhang, Weiming Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 VTFusion通过自适应特征提取和多模态融合模块,提升少样本异常检测性能,实现96.8%的AUROC。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14274 2026-01-22 cs.LG cs.AI 83%

Divide and Refine: Enhancing Multimodal Representation and Explainability for Emotion Recognition in Conversation

分割与精炼:提升对话中情感识别的多模态表示与可解释性

Anh-Tuan Mai, Cam-Van Thi Nguyen, Duc-Trong Le

机构 * VNU University of Engineering and Technology(越南工程大学) FPT Software AI Center(FPT软件人工智能中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出DnR框架,通过分割和精炼多模态表示提升对话中情感识别的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02438 2026-01-21 cs.SE cs.AI cs.CR 83%

Focus on What Matters: Fisher-Guided Adaptive Multimodal Fusion for Vulnerability Detection

聚焦关键要素:基于Fisher信息的自适应多模态融合用于漏洞检测

Yun Bian, Yi Chen, HaiQuan Wang, ShiHao Li, Zhe Cui

机构 * University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出TaCCS-DFA框架,通过Fisher信息引导的自适应多模态融合,提升漏洞检测的F1分数,同时降低推理延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05097 2026-01-14 cs.CL 83%

MultiCheck: Strengthening Web Trust with Unified Multimodal Fact Verification

MultiCheck: 通过统一多模态事实验证加强网络信任

Aditya Kishore, Gaurav Kumar, Jasabanta Patro

机构 * IISER Bhopal(比哈尔IISER)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 MultiCheck通过统一多模态事实验证框架,提升网络信任,具备高效、透明和抗噪能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05538 2026-01-12 cs.CV 83%

DIFF-MF: A Difference-Driven Channel-Spatial State Space Model for Multi-Modal Image Fusion

DIFF-MF: 一种基于差分驱动的通道-空间状态空间模型用于多模态图像融合

Yiming Sun, Zifan Ye, Qinghua Hu, Pengfei Zhu

机构 * School of Automation, Southeast University(东南大学自动化学院) Low-Altitude Intelligence Lab, Xiong’an National Innovation Center Technology Co., Ltd.(雄安国家创新中心技术有限公司低空智能实验室) Xiong’an Guochuang Lantian Technology Co., Ltd.(雄安国创蓝天科技有限公司) School of Artificial Intelligence, Tianjin University(天津大学人工智能学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 DIFF-MF通过差分驱动的通道-空间状态空间模型,有效整合多模态图像信息,提升融合图像的质量和显著性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02249 2026-01-06 cs.CV 83%

SLGNet: Synergizing Structural Priors and Language-Guided Modulation for Multimodal Object Detection

SLGNet: 结合结构先验与语言引导调制的多模态目标检测

Xiantai Xiang, Guangyao Zhou, Zixiao Wen, Wenshuai Li, Ben Niu, Feng Wang, Lijia Huang, Qiantong Wang, Yuhan Liu, Zongxu Pan, Yuxin Hu

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所) Key Laboratory of Target Cognition and Application Technology, Chinese Academy of Sciences(中国科学院目标认知与应用技术重点实验室) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 SLGNet通过结合结构先验与语言引导调制,在冻结的ViT基础上实现高效多模态目标检测,提升环境适应性和检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01339 2026-01-06 cs.CV 83%

Achieving Fine-grained Cross-modal Understanding through Brain-inspired Hierarchical Representation Learning

通过脑启发的分层表征学习实现细粒度跨模态理解

Weihang You, Hanqi Jiang, Yi Pan, Junhao Chen, Tianming Liu, Fei Dou

机构 * School of Computing, University of Georgia, Athens, GA, USA(计算学院,佐治亚大学,亚特兰大,GA,USA)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 NeuroAlign通过脑启发的分层表征学习,在fMRI-视频对齐中实现细粒度跨模态理解,提升跨模态检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00352 2026-01-05 cs.CV 83%

OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning

OmniVaT:单域泛化用于多模态视觉-触觉学习

Liuxiang Qiu, Hui Da, Yuzhen Niu, Tiesong Zhao, Yang Cao, Zheng-Jun Zha

机构 * Fujian Key Laboratory for Intelligent Processing and Wireless Transmission of Media Information(福建智能媒体信息处理与无线传输重点实验室) College of Physics and Information Engineering(物理与信息工程学院) Fuzhou University(福州市大学) College of Computer and Data Science(计算机与数据科学学院) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition(MoE脑启发智能感知与认知重点实验室) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 OmniVaT通过多模态分数傅里叶适配器和离散树生成模块,首次实现单域泛化多模态视觉-触觉学习任务,提升跨领域适应性与泛化性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24324 2026-01-01 cs.LG cs.AI 83%

Empower Low-Altitude Economy: A Reliability-Aware Dynamic Weighting Allocation for Multi-modal UAV Beam Prediction

赋能低空经济:一种可靠性感知的动态权重分配用于多模态无人机波束预测

Haojin Li, Anbang Zhang, Chen Sun, Chenyuan Feng, Kaiqian Qu, Tony Q. S. Quek, Haijun Zhang

机构 * University of Science and Technology Beijing(北京科技大学) Sony China Research Laboratory(索尼中国研究院) School of Control Science and Engineering, Shandong University(山东大学控制科学与工程学院) Southeast University(东南大学) College of Computer Science, University of Exeter(埃克塞特大学计算机学院) Information Systems Technology and Design Pillar, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出SaM2B框架,通过可靠性感知的动态权重分配和跨模态对比学习,提升多模态无人机波束预测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21508 2025-12-29 cs.CV 83%

Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification

固定预算参数高效训练结合冻结编码器提升多模态胸片分类

Md Ashik Khan, Md Nahid Siddique

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, India(计算机科学与工程系,印度理工学院Kharagpur分校) Knight Foundation School of Computing and Information Sciences, Florida International University, Florida, USA(骑士基金会计算与信息科学学院,佛罗里达国际大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本研究通过冻结编码器的参数高效训练策略,在降低计算成本的同时提升了多模态胸片分类的性能。

Comments Accepted at the 2025 28th International Conference on Computer and Information Technology (ICCIT). 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20670 2025-12-25 cs.LG cs.AI 83%

Disentangling Fact from Sentiment: A Dynamic Conflict-Consensus Framework for Multimodal Fake News Detection

区分事实与情感:一种动态冲突-共识框架用于多模态虚假新闻检测

Weilin Zhou, Zonghao Ying, Junjie Mu, Shengwei Tian, Quanchen Zou, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang

机构 * Xinjiang University(新疆大学) AI Security Lab(360AI安全实验室) Beihang University(北航) South China University of Technology(华南理工大学) Fudan University(复旦大学) Shenzhen University(深圳大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出动态冲突-共识框架,通过区分事实与情感空间,利用物理启发式特征动态和冲突-共识机制,提升多模态虚假新闻检测的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18986 2025-12-23 cs.LG cs.AI 83%

R-GenIMA: Integrating Neuroimaging and Genetics with Interpretable Multimodal AI for Alzheimer's Disease Progression

R-GenIMA:整合神经影像与基因组学的可解释多模态AI用于阿尔茨海默病进展

Kun Zhao, Siyuan Dai, Yingying Zhang, Guodong Liu, Pengfei Gu, Chenghua Lin, Paul M. Thompson, Alex Leow, Heng Huang, Lifang He, Liang Zhan, Haoteng Tang

机构 * Eli and Lilly company(艾利和利公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 R-GenIMA通过整合神经影像与基因组学,利用可解释的多模态AI方法,实现了对阿尔茨海默病进展的精准预测与机制揭示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15707 2025-12-18 cs.CV 83%

GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection

GateFusion:用于活动说话检测的分层门控跨模态融合

Yu Wang, Juhyung Ha, Frangil M. Ramirez, Yuchen Wang, David J. Crandall

机构 * Indiana University(印第安纳大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 GateFusion通过分层门控融合解码器提升活动说话检测的跨模态融合效果,实现新的SOTA结果。

Comments accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15250 2025-12-18 cs.LG cs.AI 83%

Leveraging Foundational Models and Simple Fusion for Multi-modal Physiological Signal Analysis

利用基础模型和简单融合进行多模态生理信号分析

Youssef Ghallab, Omar Iraqy, Mohamed Kandil, Mohamed Ashraf, Saadeldine Eletter, Morougue Ghazal, Ayman Khalafallah, Nagwa El-Makky

机构 * Computer and Communication Engineering Department, Alexandria University(亚历山大大学计算机与通信工程系) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出利用基础模型和简单融合方法,通过双掩码策略和对称编码器提升多模态生理信号分析的性能,实现情绪识别的高精度结果。

Comments Published at NeurIPS 2025 Workshop on Foundation Models for the Brain and Body

详情

展开后加载摘要…

URL PDF HTML 收藏