arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-03 至 2026-03-03 共收录 30 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 30 篇

2508.11999 2026-03-03 cs.CV cs.AI cs.IR cs.LG 90%

MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding

MOON: 基于生成式多模态大语言模型的电商产品理解多模态表示学习

Daoze Zhang, Chenghan Fu, Zhanheng Nie, Jianyu Liu, Wanxian Guan, Yuan Gao, Jun Song, Pengjie Wang, Jian Xu, Bo Zheng

机构 * Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 MOON通过生成式多模态大语言模型改进电商产品理解的多模态表示学习,引入引导的MoE模块、核心语义区域检测和负采样策略,提升零样本性能和泛化能力。

Comments Accepted by WSDM 2026 (oral). 11 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01536 2026-03-03 cs.IR cs.MM 88%

CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation

CLEAR: 多模态去冗余中的跨模态空域投影

Hao Zhan, Yihui Wang, Yonghui Yang, Danyang Yue, Yu Wang, Pengyang Shao, Fei Shen, Fei Liu, Le Wu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.MM

AI总结 CLEAR通过显式减少跨模态冗余提升多模态推荐性能,采用子空间投影方法增强表示学习动态。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00695 2026-03-03 cs.CV 88%

STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification

STMI: 基于跨模态超图交互的多模态目标重识别分割引导令牌调节

Xingguo Xu, Zhanyu Liu, Weixiang Zhou, Yuansheng Gao, Junjie Cao, Yuhao Wang, Jixiang Luo, Dell Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV

AI总结 STMI通过分割引导特征调节、语义令牌重新分配和跨模态超图交互,提升多模态目标重识别的性能与鲁棒性。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16479 2026-03-03 eess.IV cs.AI cs.CV 84%

Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization

解耦的多模态学习:组织学与转录组学用于癌症表征

Yupei Zhang, Xiaofei Wang, Anran Liu, Lequan Yu, Chao Li

机构 * Department of Clinical Neurosciences, University of Cambridge, UK(剑桥大学临床神经科学系) Department of Health Technology & Informatics, The Hong Kong Polytechnic University(香港理工大学健康科技与信息学系) Department of Statistics and Actuarial Science, The University of Hong Kong(香港大学统计与精算科学系) Department of Clinical Neurosciences and Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学临床神经科学系和应用数学与理论物理系;邓迪大学科学与工程学院和医学院) School of Science and Engineering and School of Medicine, University of Dundee, UK

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了解耦的多模态学习框架,通过分解组织学和转录组数据以提高癌症表征的准确性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01784 2026-03-03 cs.CR cs.AI 83%

Co-Evolutionary Multi-Modal Alignment via Structured Adversarial Evolution

基于结构对抗进化的多模态对齐

Guoxin Shi, Haoyu Wang, Zaihui Yang, Yuxing Wang, Yongzhe Chang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 本文提出CEMA框架,通过共进化对抗提升多模态对齐的鲁棒性和安全性。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00482 2026-03-03 cs.CV cs.IT math.IT 83%

TokenCom: Vision-Language Model for Multimodal and Multitask Token Communications

TokenCom: 一种用于多模态和多任务令牌通信的视觉-语言模型

Feibo Jiang, Siwei Tu, Li Dong, Xiaolong Li, Kezhi Wang, Cunhua Pan, Zhu Han, Jiangzhou Wang

机构 * Hunan Provincial Key Laboratory of Intelligent Computing and Language Information Processing, Hunan Normal University(湖南省级智能计算与语言信息处理重点实验室,湖南师范大学) School of Information Science and Engineering, Hunan Normal University(信息科学与工程学院,湖南师范大学) Changsha Social Laboratory of Artificial Intelligence, Hunan University of Technology and Business(长沙人工智能社会实验室,湖南工业大学) School of Computer Science, Hunan University of Technology and Business(计算机科学学院,湖南工业大学) Department of Computer Science, Brunel University London(伦敦布鲁内尔大学计算机科学系) National Mobile Communications Research Laboratory, Southeast University(东南大学国家移动通信研究中心) Department of Electrical and Computer Engineering, University of Houston(电子与计算机工程系,休斯顿大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 TokenCom提出了一种新的视觉-语言模型框架TaiChi,通过双视觉分词器和双向注意力网络提升多模态和多任务令牌通信的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00720 2026-03-03 cs.LG 82%

MARS: Harmonizing Multimodal Convergence via Adaptive Rank Search

MARS: 通过自适应排名搜索实现多模态收敛的统一

Minkyoung Cho, Insu Jang, Shuowei Jin, Zesen Zhao, Adityan Jothi, Ethem F. Can, Min-Hung Chen, Z. Morley Mao

机构 * University of Michigan(密歇根大学) NVIDIA(英伟达)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract)

AI总结 MARS通过自适应排名搜索优化多模态大语言模型微调,平衡训练动态并提升性能。

Comments 17 pages; Project Page: this https URL: https://minkyoungcho.github.io/mars/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02162 2026-03-03 cs.CV 79%

Bridging the gap between Performance and Interpretability: An Explainable Disentangled Multimodal Framework for Cancer Survival Prediction

弥合性能与可解释性之间的鸿沟:一种可解释的解耦多模态框架用于癌症生存预测

Aniek Eijpe, Soufyan Lakbir, Melis Erdal Cesur, Sara P. Oliveira, Angelos Chatzimparmpas, Sanne Abeln, Wilson Silva

机构 * AI Technology for Life(人工智能技术与生命科学) Department of Information and Computing Sciences(信息与计算科学系) Department of Biology(生物学系) Utrecht University(乌得勒支大学) Department of Metabolic Diseases(代谢疾病部门) Wilhelmina Children’s Hospital(维廉明娜儿童医院) University Medical Center Utrecht(乌得勒支大学医学中心) Regenerative Medicine Center Utrecht(乌得勒支再生医学中心) Computational Pathology(计算病理学) Department of Pathology(病理学系) The Netherlands Cancer Institute(荷兰癌症研究所) Visualization and Graphics(可视化与图形学) The Netherlands Cancer Institute, Amsterdam(荷兰癌症研究所,阿姆斯特丹)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 DIMAFx通过解耦多模态表示提升癌症生存预测的性能与可解释性,揭示了多模态交互和生物学信息。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01758 2026-03-03 cs.CV 79%

Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining

通过语言枢轴预训练统一异构多模态遥感检测

Yuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng, Xiang Li, Jian Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 BabelRS通过语言枢轴预训练框架统一异构多模态遥感检测,解耦模态对齐与任务学习,提升训练稳定性与检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09285 2026-03-03 cs.CV 79%

Spotlight on Token Perception for Multimodal Reinforcement Learning

多模态强化学习中的token感知聚焦

Siyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo, Zefeng He, Daizong Liu, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Nanjing University(南京大学) Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出VPPO算法,通过token感知优化提升多模态强化学习的视觉推理能力。

Comments Accepted by ICLR 2026, project page: https://github.com/huaixuheqing/VPPO-RL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03214 2026-03-03 cs.CV 79%

RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion

RTGMFF:基于ROI驱动文本生成和多模态特征融合的增强型fMRI脑部疾病诊断

Junhao Jia, Yifei Sun, Yunyou Liu, Cheng Yang, Changmiao Wang, Feiwei Qin, Yong Peng, Wenwen Min

机构 * Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University(浙江大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) Yunnan University(云南大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 RTGMFF通过结合ROI驱动文本生成和多模态特征融合,提升fMRI在脑部疾病诊断中的准确性。

Comments The paper has been accepted by BIBM 2025

Journal ref 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE, 2025, pp. 2301-2308

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22283 2026-03-03 cs.CV 79%

Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment

重新审视在跨模态不匹配下的LVLMs视觉标记减少

Rui Xu, Yunke Wang, Yong Luo, Bo Du

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出VisionDrop方法,通过视觉-only修剪框架减少LVLMs中的视觉标记,无需额外训练,提升推理效率并保持性能。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00157 2026-03-03 cs.CV 79%

FujiView: Multimodal Late-Fusion for Predicting Scenic Visibility

FujiView: 多模态晚期融合用于预测风景可见性

Bryceton Bible, Shah Md Nehal Hasnaeen, Hairong Qi

机构 * University of Tennessee, Knoxville(田纳西大学,科文克顿)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 FujiView通过融合摄像头图像与气象数据,实现风景可见性的多模态预测,展示了在短期和长期预测中的不同方法效果。

Comments 9 pages (including references), 8 figures, 2 tables. Accepted to the IEEE/CVF WACV 2026 proceedings. Introduces a large human-labeled Mount Fuji visibility dataset; public release forthcoming

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00699 2026-03-03 astro-ph.IM 78%

Deep learning-based astronomical multimodal data fusion: A comprehensive review

基于深度学习的天文多模态数据融合:综述

Wujun Shao, Dongwei Fan, Chenzhou Cui, Yunfei Xu, Shirui Wei, Xin Lyu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文综述了基于深度学习的天文多模态数据融合方法,探讨了其在数据融合中的应用、挑战及未来发展方向。

Journal ref Information Fusion 130 (2026) 104103

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00193 2026-03-03 q-bio.QM 78%

Multimodal Alignment Improves Generalizability of Genomic Biomarker Prediction in Computational Pathology

多模态对齐提升了计算病理学中基因组生物标志物预测的泛化能力

Ekaterina Redekop, Eric Zimmermann, Ava P Amini, Alex X Lu, Neil Tenenholtz, James Brian Hall, Lorin Crawford, Kristen A Severson

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 MARBLE通过多模态对比预训练策略,将组织病理学图像与基因组生物标志物的表示对齐,提升计算病理学中基因组生物标志物预测的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01210 2026-03-03 cs.CV cs.RO 74%

OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation

OmniVLA:具有统一多传感器感知的物理基础多模态VLA

Heyu Guo, Shanmu Wang, Ruichun Ma, Shiqi Jiang, Yasaman Ghasempour, Omid Abari, Baining Guo, Lili Qiu

机构 * Princeton University(普林斯顿大学) University of California, Los Angeles(加州大学洛杉矶分校) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 OmniVLA通过整合多种传感器模态,提升机器人操作的感知能力与任务成功率。

Comments Accepted by ICRA'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01720 2026-03-03 cs.CV 70%

Preoperative-to-intraoperative Liver Registration for Laparoscopic Surgery via Latent-Grounded Correspondence Constraints

腹腔手术中基于潜在证据的预手术到手术过程肝脏注册

Ruize Cui, Jialun Pei, Haiqiao Wang, Jun Zhou, Jeremy Yuen-Chun Teoh, Pheng-Ann Heng, Jing Qin

机构 * The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学) The Chinese University of Hong Kong, Hong Kong, China(香港中文大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出Land-Reg框架,通过显式学习潜在证据的2D-3D地标对应关系,提升腹腔手术中预手术到术中肝脏的跨模态注册精度与可解释性。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02948 2026-03-03 cs.AI cs.LG physics.ao-ph 70%

FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead

风梧:将高技能全球中期天气预报推至10天提前期

Kang Chen, Tao Han, Junchao Gong, Lei Bai, Fenghua Ling, Jing-Jia Luo, Xi Chen, Leiming Ma, Tianning Zhang, Rui Su, Yuanzheng Ci, Bin Li, Xiaokang Yang, Wanli Ouyang

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 FengWu通过多模态和多任务框架提升全球中期天气预报能力,首次实现10.75天提前期的高精度预测。

Comments 12 pages

Journal ref Commun. Earth Environ. 6, 518 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00409 2026-03-03 cs.CV 70%

SSR: Pushing the Limit of Spatial Intelligence with Structured Scene Reasoning

SSR:通过结构化场景推理推动空间智能的极限

Yi Zhang, Youya Xia, Yong Wang, Meng Song, Xin Wu, Wenjun Wan, Bingbing Liu, AiXue Ye, Hongbo Zhang, Feng Wen

机构 * Foundation Model Department, Huawei(华为基础模型部门)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 SSR通过结构化场景推理框架,在减少对齐成本的同时,实现了高效的空间智能,优于更大规模模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00046 2026-03-03 cs.LG cs.AI 70%

REMIND: Rethinking Medical High-Modality Learning under Missingness--A Long-Tailed Distribution Perspective

REMIND: 重新思考医疗多模态学习中的缺失性——从长尾分布视角

Chenwei Wu, Zitao Shuai, Liyue Shen

机构 * University of Michigan(密歇根大学)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

AI总结 REMIND从长尾分布视角重新思考医疗多模态学习中的高模态缺失问题,提出组专用混合专家架构和分布鲁棒优化策略,有效提升尾部模态组合的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01725 2026-03-03 cs.CV 57%

Learning Domain-Aware Task Prompt Representations for Multi-Domain All-in-One Image Restoration

学习多领域任务提示表示以实现多领域一体化图像修复

Guanglu Dong, Chunlei Li, Chao Ren, Jingliang Hu, Yilei Shi, Xiao Xiang Zhu, Lichao Mou

机构 * Sichuan University(四川大学) MedAI Technology (Wuxi) Co. Ltd.(MedAI技术(无锡)有限公司) Technical University of Munich(慕尼黑技术大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 提出DATPRL-IR,通过多领域任务提示表示学习实现多领域一体化图像修复,优于现有方法并具备强泛化能力

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01135 2026-03-03 cs.AI 57%

FCN-LLM: Empower LLM for Brain Functional Connectivity Network Understanding via Graph-level Multi-task Instruction Tuning

FCN-LLM: 通过图级多任务指令微调赋能LLM理解脑功能连接网络

Xingcan Hu, Wei Wang, Li Xiao

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 FCN-LLM通过图级多任务指令微调使LLM理解脑功能连接网络,提升其在临床任务中的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02742 2026-03-03 cs.LG cs.AI 57%

Entropy-Guided Dynamic Tokens for Graph-LLM Alignment in Molecular Understanding

熵引导的动态令牌用于分子理解中的图-语言模型对齐

Zihao Jing, Qiuhao Zeng, Ruiyi Fang, Yan Sun, Boyu Wang, Pingzhao Hu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 EDT-Former通过熵引导的动态令牌生成,在无需微调LLM主干的情况下实现图编码器与LLM的对齐,提升分子图理解的效率和泛化能力。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16953 2026-03-03 cs.CV 57%

Towards Real Zero-Shot Camouflaged Object Segmentation without Camouflaged Annotations

面向无遮蔽标注的零样本遮蔽物分割

Cheng Lei, Jie Fan, Xinran Li, Tianzhu Xiang, Ao Li, Ce Zhu, Le Zhang

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Space42

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种无需遮蔽标注的零样本遮蔽物分割框架,通过结合MIM、M-LLM和MFA机制,实现高效分割与快速推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00793 2026-03-03 cs.CV 57%

Neural Functional Alignment Space: Brain-Referenced Representation of Artificial Neural Networks

神经功能对齐空间:人工神经网络的脑参考表示

Ruiyu Yan, Hanqi Jiang, Yi Pan, Xiaobo Li, Tianming Liu, Xi Jiang, Lin Zhao

机构 * Tandon School of Engineering, New York University(纽约大学工程学院) School of Computing, University of Georgia(佐治亚大学计算机学院) Department of Biomedical Engineering, New Jersey Institute of Technology(新泽西理工学院生物医学工程系) The Clinical Hospital of Chengdu Brain Science Institute, MOE Key Laboratory for NeuroInformation,School of Life Science and Technology, University of Electronic Science and Technology of China(成都脑科学研究所临床医院,国家神经信息重点实验室,电子科技大学生命科学与技术学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出神经功能对齐空间,通过建模神经网络的动态表示,揭示了在脑参考空间中的结构化组织,包括模态特定聚类和跨模态收敛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00609 2026-03-03 cs.CV 57%

Linking Modality Isolation in Heterogeneous Collaborative Perception

异构协作感知中的模态隔离问题

Changxing Liu, Zichen Chao, Siheng Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 CodeAlign通过跨模态特征-代码-特征翻译有效解决异构协作感知中的模态隔离问题,显著降低训练参数和通信负载,提升感知性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19534 2026-03-03 cs.RO cs.AI 57%

Large Language Model-Assisted UAV Operations and Communications: A Multifaceted Survey and Tutorial

大型语言模型辅助的无人机操作与通信:多方面的综述与教程

Yousef Emami, Hao Zhou, Radha Reddy, Atefeh Hajijamali Arani, Biliang Wang, Kai Li, Luis Almeida, Zhu Han

机构 * IEEE

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本文综述了大型语言模型在无人机操作与通信中的应用,探讨了LLMs在提升UAV智能方面的多方面技术与未来研究方向。

Comments 40 pages, 10 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15663 2026-03-03 cs.CV 57%

MSSPlace: Multi-Sensor Place Recognition with Visual and Text Semantics

MSSPlace: 多传感器位置识别与视觉和文本语义

Alexander Melekhin, Dmitry Yudin, Ilia Petryashin, Vitaly Bezuglyj

机构 * Intelligent Transport Laboratory, Moscow Institute of Physics and Technology(智能交通实验室,莫斯科物理技术学院) Artificial Intelligence Research Institute (AIRI)(人工智能研究机构(AIRI))

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 MSSPlace通过整合多传感器数据和视觉文本语义,提升位置识别性能,达到最先进的效果。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00156 2026-03-03 cs.CV 57%

BiCLIP: Bidirectional and Consistent Language-Image Processing for Robust Medical Image Segmentation

BiCLIP: 用于鲁棒医学图像分割的双向和一致语言-图像处理

Saivan Talaei, Fatemeh Daneshfar, Abdulhady Abas Abdullah, Mustaqeem Khan

机构 * Department of Computer Engineering, University of Kurdistan, Iran(伊朗库尔德大学计算机工程系) Artificial Intelligence and Innovation Centre, University of Kurdistan, Erbil, Iraq(伊拉克埃尔比尔库尔德大学人工智能与创新中心) College of Information Technology, United Arab Emirates University, UAE(阿联酋大学信息科技学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 BiCLIP通过双向多模态融合和一致性目标提升医学图像分割的鲁棒性,有效应对标注稀少和临床伪影挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21028 2026-03-03 cs.LG 50%

TRIDENT: Tri-Modal Molecular Representation Learning with Taxonomic Annotations and Local Correspondence

TRIDENT:结合分类注释和局部对应关系的三模态分子表示学习

Feng Jiang, Mangal Prakash, Hehuan Ma, Jianyuan Deng, Yuzhi Guo, Amina Mollaysa, Tommaso Mansi, Rui Liao, Junzhou Huang

机构 * University of Texas at Arlington(德克萨斯大学阿灵顿分校) Johnson & Johnson Innovative Medicine(强生创新医学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 TRIDENT通过整合SMILES、文本和分类功能注释,学习丰富的分子表示,从而在多个下游任务中取得最佳性能。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏