arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2601.10092 2026-01-16 cs.LG cs.AI 79%

LeMoF: Level-guided Multimodal Fusion for Heterogeneous Clinical Data

LeMoF:面向异构临床数据的层级引导多模态融合

Jongseok Kim, Seongae Kang, Jonghwan Shin, Yuhan Lee, Ohyun Jo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 LeMoF通过层级引导的多模态融合方法,在异构临床数据中提升预测稳定性与判别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08503 2026-01-14 cs.LG cs.AI 79%

Temporal Fusion Nexus: A task-agnostic multi-modal embedding model for clinical narratives and irregular time series in post-kidney transplant care

时间融合 nexus:一种任务无关的多模态嵌入模型,用于术后肾移植护理中的临床叙述和不规则时间序列

Aditya Kumar, Simon Rauch, Mario Cypko, Marcel Naik, Matthieu-P Schapranow, Aadil Rashid, Fabian Halleck, Bilgin Osmanodja, Roland Roller, Lars Pape, Klemens Budde, Mario Schiffer, Oliver Amft

机构 * Hahn-Schickard(哈恩-施克特德研究所) University of Freiburg(弗赖堡大学) Charité University Medical Center(查理医院大学医学中心) Hasso Plattner Institute for Digital Engineering, University of Potsdam(哈索·platner数字工程研究所,波茨坦大学) DFKI(德意志联邦防务研究院) University Hospital Essen(埃森大学医院) University Hospital Erlangen(埃尔兰根大学医院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 TFN是一种多模态嵌入模型,通过整合临床文本和不规则时间序列数据,在术后肾移植护理中提高了移植物丢失、排斥和死亡预测的性能。

Comments 31 pages, 9 figures, 3 tables. A supplementary file is also available

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07856 2026-01-14 quant-ph cs.AI cs.LG 79%

Feature Entanglement-based Quantum Multimodal Fusion Neural Network

基于特征纠缠的量子多模态融合神经网络

Yu Wu, Qianli Zhou, Jie Geng, Xinyang Deng, Wen Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出基于特征纠缠的量子多模态融合神经网络,通过量子计算框架解决多模态学习中的精度、可解释性和复杂性矛盾,实现高效且可解释的多模态融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01939 2026-01-13 cs.AI 79%

OpenSocInt: A Multi-modal Training Environment for Human-Aware Social Navigation

OpenSocInt: 一种多模态训练环境用于人感知的社会导航

Victor Sanchez, Chris Reinke, Ahamed Mohamed, Xavier Alameda-Pineda

机构 * Inria at Univ. Grenoble Alpes, LJK, CNRS(Inria与格勒诺布尔阿尔卑斯大学、LJK、CNRS)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 OpenSocInt是一个开源多模态训练环境,用于研究人感知的社会导航,支持不同感知特征的编码与融合以及多种代理的训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00521 2026-01-13 cs.LG cs.CL q-bio.QM 79%

Rep3Net: An Approach Exploiting Multimodal Representation for Molecular Bioactivity Prediction

Rep3Net:一种利用多模态表示进行分子生物活性预测的方法

Sabrina Islam, Md. Atiqur Rahman, Md. Bakhtiar Hasan, Md. Hasanul Kabir

机构 * Computer Science and Engineering, Islamic University of Technology, Bangladesh(计算机科学与工程,伊斯兰技术大学,孟加拉国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 Rep3Net通过融合分子描述符、图特征和SMILES嵌入,提升了PARP1生物活性预测的准确性,并在药物发现中展示了高效能的多模态方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06848 2026-01-13 cs.CL 79%

Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model

可解释的多模态基于方面的情感分析与依赖引导的大语言模型

Zhongzheng Wang, Yuanhe Tian, Hongzhi Wang, Yan Song

机构 * Harbin Institute of Technology(哈尔滨工程大学) Zhongguancun Academy(中关村学院) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出了一种基于依赖引导的大语言模型的多模态情感分析方法,通过生成可解释的自然语言解释来提升情感分类的准确性。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12195 2026-01-13 cs.CL 79%

Browse and Concentrate: Comprehending Multimodal Content via prior-LLM Context Fusion

浏览与聚焦:通过先验LLM上下文融合理解多模态内容

Ziyue Wang, Chi Chen, Yiqi Zhu, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu

机构 * Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China(计算机科学与技术系,人工智能研究院,清华大学,北京,中国) Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China(人工智能产业研究院(AIR),清华大学,北京,中国) Institute of Intelligent Computing, Alibaba Group(智能计算研究院,阿里巴巴集团) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国) Jiangsu Collaborative Innovation Center for Language Competence, Jiangsu, China(江苏省语言能力协同创新中心,江苏,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出浏览与聚焦两阶段范式,通过融合先验LLM上下文提升多模态内容理解,显著提升多图像场景的性能。

Comments 17 pages, 5 figures

Journal ref ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14329 2026-01-12 cs.MM 79%

TF-Mamba: Text-enhanced Fusion Mamba with Missing Modalities for Robust Multimodal Sentiment Analysis

具有缺失模态的文本增强融合Mamba用于鲁棒多模态情感分析

Xiang Li, Xianfu Cheng, Dezhuang Miao, Xiaoming Zhang, Zhoujun Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 TF-Mamba通过文本增强融合框架,有效处理多模态情感分析中缺失模态的问题,提升模型鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04736 2026-01-09 cs.CL 79%

AM$^3$Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs

AM$^3$Safety: 向多模态多轮安全对齐的数据高效方法

Han Zhu, Jiale Chen, Chengkun Cai, Shengjie Sun, Haoran Li, Yujin Zhou, Chi-Min Chan, Pengcheng Wen, Lei Li, Sirui Han, Yike Guo

机构 * Hong Kong University of Science and Technology(香港科技大学) Zhongshan School of Medicine, SUN YAT-SEN UNIVERSITY(中山医学院,孙中山大学) University of Edinburgh(爱丁堡大学) University of Washington(华盛顿大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

AI总结 AM$^3$Safety通过结合冷启动拒绝阶段和组相对策略优化,有效提升多模态多轮对话的安全性,降低攻击成功率并增强模型的无害与帮助维度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03296 2026-01-09 cs.CL cs.LG 79%

Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling

通过政策对齐推理和分层标注实现可信的多模态审核

Anqi Li, Wenwei Jin, Jintao Tong, Pengda Qin, Weijia Li, Guo Lu

机构 * Shanghai Jiao Tong University(上海交通大学) Xiaohongshu Inc.(小红书公司) Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 Hi-Guard通过政策对齐推理和分层标注提升多模态审核的准确性、泛化性和可解释性。

Comments Accepted by KDD 2026. Code is available at https://github.com/lianqi1008/Hi-Guard

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21582 2026-01-06 cs.CV 79%

Data-Augmented Multimodal Feature Fusion for Multiclass Visual Recognition of Oral Cancer Lesions

数据增强多模态特征融合用于口腔癌病变的多类视觉识别

Joy Naoum, Revana Salama, Ali Hamdi

机构 * MSA University(MSA大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出一种数据增强驱动的多模态特征融合框架,用于提升口腔癌病变的多类视觉识别性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00273 2026-01-05 eess.IV cs.CV 79%

UKAN-EP: Enhancing U-KAN with Efficient Attention and Pyramid Aggregation for 3D Multi-Modal MRI Brain Tumor Segmentation

UKAN-EP: 通过高效的注意力和金字塔聚合增强U-KAN以实现3D多模态MRI脑肿瘤分割

Yanbing Chen, Tianze Tang, Taehyo Kim, Hai Shu

机构 * Department of Biostatistics, School of Global Public Health, New York University(生物统计学系,全球公共卫生学院,纽约大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 UKAN-EP通过引入高效的通道注意力和金字塔特征聚合模块,提升3D多模态MRI脑肿瘤分割的准确性和效率。

Journal ref BMC Medical Imaging, Volume 25, article number 517, (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22748 2026-01-01 cs.CV 79%

TrimTokenator-LC: Towards Adaptive Visual Token Pruning for Large Multimodal Models with Long Contexts

TrimTokenator-LC: 向大型多模态模型长上下文的自适应视觉标记修剪迈进

Hao Zhang, Mengsi Lyu, Bo Huang, Yulong Ao, Yonghua Lin

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 TrimTokenator-LC通过自适应视觉标记修剪方法,在长上下文和多图像场景中有效减少视觉标记数量,同时保持性能。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22503 2025-12-30 cs.CV 79%

SCAFusion: A Multimodal 3D Detection Framework for Small Object Detection in Lunar Surface Exploration

SCAFusion: 一种针对月球表面探索的小目标多模态3D检测框架

Xin Chen, Kang Luo, Yangyi Xiao, Hesheng Wang

机构 * Department of Automation, Key Laboratory of System Control and Information Processing of Ministry of Education, Key Laboratory of Marine Intelligent Equipment and System of Ministry of Education, Shanghai Engineering Research Center of Intelligent Control and Management, Shanghai Jiao Tong University(自动化系、教育部系统控制与信息处理重点实验室、教育部海洋智能装备与系统重点实验室、上海智能控制与管理工程研究中心、上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 SCAFusion提出了一种针对月球表面探索的小目标多模态3D检测框架,通过改进的特征对齐和坐标注意机制提升小目标检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21916 2025-12-29 cs.CV 79%

Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition

补丁作为节点:面向多模态动作识别的人本图表示学习

Zeyu Liang, Hailun Xia, Naichuan Zheng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出PAN框架,通过人本图表示学习提升多模态动作识别性能,结合双路径和统一网络结构实现高效融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21897 2025-12-29 cs.LG cs.AI 79%

MMCTOP: A Multimodal Textualization and Mixture-of-Experts Framework for Clinical Trial Outcome Prediction

MMCTOP: 一种用于临床试验结果预测的多模态文本化与专家混合框架

Carolina Aparício, Qi Shi, Bo Wen, Tesfaye Yadete, Qiwei Han

机构 * Nova School of Business and Economics(诺瓦商学院) Hogarthian Technologies(霍加斯技术) School of Medicine(医学院) Oregon Health & Science University(俄勒冈健康与科学大学) Cleveland Clinic(克利夫兰诊所) IBM Research(IBM研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 MMCTOP通过多模态文本化与专家混合框架,提升临床试验结果预测的精度与稳定性。

Comments 15 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19443 2025-12-29 cs.CV 79%

D2Pruner: Debiased Importance and Structural Diversity for MLLM Token Pruning

D2Pruner: 用于MLLM标记剪枝的去偏重要与结构多样性

Evelyn Zhang, Fufu Yu, Aoqi Wu, Zichen Wen, Ke Yan, Shouhong Ding, Biqing Qi, Linfeng Zhang

机构 * Tencent YouTu Lab(腾讯YouTu实验室)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 D2Pruner通过结合去偏重要与结构剪枝机制,有效提升MLLM标记剪枝的效率和保真度,尤其在细粒度定位任务中表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09626 2025-12-29 cs.SI cs.AI cs.LG 79%

Certainly Bot Or Not? Trustworthy Social Bot Detection via Robust Multi-Modal Neural Processes

确定是机器人还是不是?通过鲁棒多模态神经过程进行可信的社交机器人检测

Qi Wu, Yingguang Yang, hao liu, Hao Peng, Buyun He, Yutong Xia, Yong Liao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 本研究提出鲁棒多模态神经过程框架,通过增强多模态神经过程的鲁棒性来检测社交机器人,同时提升不确定性估计能力。

Comments We withdraw this paper due to an error identified in the experimental setup. Specifically, the evaluation protocol described in Section 4 does not correctly reflect the intended experimental design, which may affect the validity of the reported results. To avoid potential misunderstanding by readers, we choose to withdraw this version and revise the work before resubmission

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20146 2025-12-29 cs.CV 79%

AlignFreeNet: Is Cross-Modal Pre-Alignment Necessary? An End-to-End Alignment-Free Lightweight Network for Visible-Infrared Object Detection

AlignFreeNet: 跨模态预对齐是否必要?一种端到端无对齐的轻量级网络用于可见-红外目标检测

Dingkun Zhu, Haote Zhang, Lipeng Gu, Wuzhou Quan, Fu Lee Wang, Honghui Fan, Jiali Tang, Haoran Xie, Xiaoping Zhang, Mingqiang Wei

机构 * School of Computer Science, Jiangsu University of Technology(江苏科技大学计算机科学学院) School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) School of Science and Technology, Hong Kong Metropolitan University(香港都会大学科技学院) School of Data Science, Lingnan University(岭南大学数据科学学院) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 AlignFreeNet通过无对齐融合范式,提出VCC和FCF模块,有效缓解可见-红外目标检测中的跨模态错位问题,实现端到端轻量级网络的高鲁棒性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20084 2025-12-24 cs.LG cs.AI 79%

QE-Catalytic: A Graph-Language Multimodal Base Model for Relaxed-Energy Prediction in Catalytic Adsorption

QE-Catalytic: 一种图-语言多模态基础模型,用于催化吸附中放松能量的预测

Yanjie Li, Jian Xu, Xueqing Chen, Lina Yu, Shiming Xiang, Weijun Li, Cheng-lin Liu

机构 * AnnLab(安实验室) Institute of Semiconductors, Chinese Academy of Sciences(半导体研究所,中国科学院) Zhongguancun Academy(中关村学院) State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Computer Network Information Center(计算机网络信息中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 QE-Catalytic结合语言模型与图Transformer,实现高精度催化吸附能量预测及逆向设计

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20026 2025-12-24 cs.CV 79%

MAPI-GNN: Multi-Activation Plane Interaction Graph Neural Network for Multimodal Medical Diagnosis

MAPI-GNN:多激活平面交互图神经网络用于多模态医学诊断

Ziwei Qin, Xuhui Song, Deqing Huang, Na Qin, Jun Li

机构 * Ziwei Qin(独立研究者) Xuhui Song(独立研究者) Deqing Huang(独立研究者) Na Qin(独立研究者) Jun Li(独立研究者)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 MAPI-GNN通过多激活平面交互机制,有效建模患者特异性病理关系,提升多模态医学诊断的准确性。

Comments Accepted by Proceedings of the AAAI Conference on Artificial Intelligence 40 (AAAI-26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19213 2025-12-23 cs.CV 79%

InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training

InvCoSS: 基于反向驱动的医学多模态图像预训练中的连续自监督学习

Zihao Luo, Shaohao Rui, Zhenyu Tang, Guotai Wang, Xiaosong Wang

机构 * University of Electronic Science and Technology of China(电子科技大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Brain-Computer Interface & Brain-Inspired Intelligence Key Laboratory of Sichuan Province(四川省脑机接口与脑启发智能重点实验室)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 InvCoSS通过反向生成合成图像和多尺度融合网络,实现连续自监督学习,减少存储需求并保护数据隐私。

Comments 16 pages, 10 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14972 2025-12-23 cs.CL 79%

Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies

多模态文化安全:评估框架与对齐策略

Haoyi Qiu, Kung-Hsiang Huang, Ruichen Zheng, Jiao Sun, Nanyun Peng

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Salesforce AI Research(Salesforce人工智能研究) Google DeepMind(谷歌DeepMind)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出CROSS基准和CROSS-Eval框架,评估多模态模型的文化安全能力,发现提升推理能力可改善文化对齐,但需结合监督微调和偏好微调策略以增强文化合规性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11361 2025-12-23 cs.CL 79%

VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment

VLDBench:评估具有监管对齐的多模态虚假信息

Shaina Raza, Ashmal Vayani, Aditya Jain, Aravind Narayanan, Vahid Reza Khazaie, Syed Raza Bashir, Elham Dolatabadi, Gias Uddin, Christos Emmanouilidis, Rizwan Qureshi, Mubarak Shah

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 VLDBench 是首个多模态虚假信息检测基准,通过大规模标注数据提升检测准确率,支持 AI 管治框架下的可信虚假信息分析。

Comments Accepted in Information Fusion Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17227 2025-12-22 cs.CV 79%

Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning

学习何时观察:一种解耦的课程学习框架,用于多模态推理中的战略感知

Siqi Yang, Zilve Gao, Haibo Qiu, Fanfan Liu, Peng Shi, Zhixiong Zeng, Qingmin Liao, Lin Ma

机构 * Meituan(美团) Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种解耦课程学习框架,通过分阶段训练提升多模态推理中的抽象推理与战略视觉感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20494 2025-12-22 cs.LG cs.MM 79%

Multimodal Representation Learning and Fusion

多模态表示学习与融合

Qihang Jin, Enze Ge, Yuhang Xie, Hongying Luo, Junhao Song, Ziqian Bi, Chia Xin Liang, Jibin Guan, Joe Yeong, Xinyuan Song, Junfeng Hao

机构 * AI Agent Lab, Vokram Group, United Kingdom(AI代理实验室,Vokram集团,英国) University of Bologna, Italy(博洛尼亚大学,意大利) University of Minnesota, United States(明尼苏达大学,美国) Singapore General Hospital, Singapore(新加坡中央医院,新加坡) Emory University, United States(埃默里大学,美国)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.MM

AI总结 多模态学习通过融合多种数据模态提升AI系统的理解和决策能力,旨在解决数据格式差异、输入缺失及对抗攻击等问题,推动计算机视觉、自然语言处理等领域的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11160 2025-12-17 cs.CV 79%

A Unified Framework with Multimodal Fine-tuning for Remote Sensing Semantic Segmentation

多模态微调的统一框架用于遥感语义分割

Xianping Ma, Xiaokang Zhang, Man-On Pun, Bo Huang

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)科学与工程学院) School of Information Science and Engineering, Wuhan University of Science and Technology(武汉科技大学信息科学与工程学院) Department of Geography, The University of Hong Kong(香港大学地理系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种多模态微调的统一框架,用于改进遥感语义分割的性能,通过引入新的MFNet和DFM模块,显著提升了多模态数据的分割效果。

Comments 15 pages, 11 figures

Journal ref IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1-15, 2025, Art no. 5405015

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12657 2025-12-16 cs.CV 79%

Cross-modal Fundus Image Registration under Large FoV Disparity

跨模态视网膜图像在大视野差异下的配准

Hongyang Li, Junyi Tao, Qijie Wei, Ningzhi Yang, Meng Wang, Weihong Yu, Xirong Li

机构 * Renmin University of China, Beijing, China(中国人民大学) Peking Union Medical College Hospital, Beijing, China(北京友谊医院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出CARe方法,用于解决大视野差异下的跨模态视网膜图像配准问题,通过裁剪和双拟合对齐改进配准效果。

Comments Accepted as a regular paper at MMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20991 2025-12-16 cs.CV 79%

MR-COSMO: Visual-Text Memory Recall and Direct CrOSs-MOdal Alignment Method for Query-Driven 3D Segmentation

MR-COSMO:一种用于查询驱动3D分割的视觉-文本记忆召回与直接跨模态对齐方法

Chade Li, Pengju Zhang, Yihong Wu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 MR-COSMO通过视觉-文本记忆召回与直接跨模态对齐方法,在查询驱动的3D分割中实现几何与语义特征的精确融合,提升点云分割性能。

Comments Accepted by AAAI 2026. Copyright (c) 2026, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04519 2025-12-12 cs.CV 79%

l0-Regularized Sparse Coding-based Interpretable Network for Multi-Modal Image Fusion

基于l0正则化稀疏编码的可解释网络用于多模态图像融合

Gargi Panda, Soumitra Kundu, Saumik Bhattacharya, Aurobinda Routray

机构 * Department of EE, IIT Kharagpur, India(印度IIT Kharagpur电子工程系) Rekhi Centre of Excellence for the Science of Happiness, IIT Kharagpur, India(印度IIT Kharagpur幸福科学卓越中心) Department of E&ECE, IIT Kharagpur, India(印度IIT Kharagpur电子工程与电子学系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出基于l0正则化稀疏编码的可解释网络FNet,用于多模态图像融合,通过分离独特和共同特征提升融合质量并增强下游任务性能。

Comments Accetped by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏