SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
SmartCLIP: 基于识别保证的模块化视觉-语言对齐
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; The University of Sydney(悉尼大学)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
AI总结 本文提出SmartCLIP,通过理论条件实现文本与视觉表征的灵活对齐,确保语义信息完整保留并解耦视觉表征,提升多任务性能。
Comments CVPR2025