IMAGINATOR: Pre-Trained Image+Text Joint Embeddings using Word-Level Grounding of Images
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments CVPR 2023. Code and pretrained models are available at https://github.com/google-research/big_vision/blob/main/big_vision/configs/proj/clippo/README.md
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence
专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL
Comments 7 pages
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments Published in EMNLP 2022
专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV
Comments 19 pages, 13 tables, 6 figures, Oral at NeuRIPS 2022
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2023
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.MM
Comments Submitted to APPLIED SCIENCES
专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV
Comments Accepted by SIGIR 2022
专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV
Comments CVPR 2022; code: https://github.com/uta-smile/TCL
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 eess.AS
Comments Copyright 2022 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV
Comments 20 pages, 7 figures, 6 tables, to appear in BMVC2021
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
专题命中 跨模态检索 :multi-modal(abstract);image-text(abstract);分类 cs.CV
Comments CVPR 2021 camera-ready (oral). The new version fixes a few typos and updates citations
专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments 8 pages
专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments Accepted for presentation at ICCV. Project Page: https://mwray.github.io/FGAR
专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments ECCV MULA Workshop 2018
专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments 10 pages, 8 figures, 3 tables
专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments ICCV 2017
QA-Dragon:面向知识密集型视觉问答的查询感知动态RAG系统
专题命中 跨模态检索 :multimodal(abstract,journal_ref);分类 cs.CV、cs.CL、cs.AI
AI总结 QA-Dragon通过引入领域路由器和搜索路由器,实现多模态、多轮和多跳推理,提升复杂视觉问答任务的推理性能,实验显示其在单源、多源和多轮任务中均优于基线模型。
Comments The source code for our system is released in https://github.com/jzzzzh/QA-Dragon
Journal ref 2025 KDD Cup Workshop for Multimodal Retrieval Augmented Generation
EgoCITE:面向长时程自我中心记忆的上下文增强索引与时序感知检索
机构 * University of Michigan(密歇根大学)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本研究针对长时程自我中心记忆系统的索引不可靠、忽略时序意图的问题,提出EgoCITE框架,经多数据集评估,其准确率优于基线且成本显著低于长上下文LLM智能体。
二次审视:多模态大语言模型中的免训练证据高亮
机构 * University of Modena and Reggio Emilia(摩德纳和雷焦艾米利亚大学)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本文提出Look Twice框架,通过模型注意力模式识别相关视觉区域和文本证据,提升预训练MLLMs在多模态证据利用中的表现,实验显示在多个VQA基准上效果显著。
Comments Project Page: https://aimagelab.github.io/LoT/
PHA-Net:用于文本-视频检索的基于原型的分层对齐网络
专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract)
AI总结 针对文本-视频检索中跨模态语义不匹配及计算成本高的问题,提出PHA-Net,以模态共享原型为桥梁,结合原型支持的令牌合并模块与原型对比损失,在四个基准数据集上取得显著性能提升。
MMEB-V3: 多模态嵌入模型性能差距的测量
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract)
AI总结 本文提出MMEB-V3基准,评估文本、图像、视频、音频及基于代理的多模态嵌入,发现现有模型在跨模态检索中存在显著偏差和不足。
Comments Accepted at COLM 2026
DialogueVPR:迈向对话式视觉场所识别
机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) ; Shandong Computer Science Center, Qilu University of Technology(山东省计算中心(国家超级计算济南中心),齐鲁工业大学) ; Macquarie University(麦考瑞大学) ; Spatialtemporal AI(时空人工智能公司) ; University of Macau(澳门大学) ; Aerospace Information Research Institute(航天信息研究所)
专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 研究针对语言引导地理定位中现有方法不足,提出DlgPR,将场所识别转为对话驱动推理过程。构建DlgQuest-Cities基准及统一推理框架,用课程训练DQ-pilot,通过特定指标指导学习并实验,该方法显著优于基线。
Comments Accepted to CVPR 2026
有备无患:当非序列嵌入变成异常检测器
机构 * Orange Research(Orange研究院)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS
AI总结 本文深入分析非序列多模态句子级嵌入(SONAR模型),发现某些嵌入维度对扰动敏感,可作为解码异常指标,并利用编解码一致性构建准确检测器,同时探索修正异常维度。
Comments Accepted for presentation at LREC 2026
低成本基于概念的可解释性:无训练方法能走多远?
机构 * Dept. of Computer Science and Artificial Intelligence, University of Granada (UGR)(计算机科学与人工智能系,格拉纳达大学(UGR))
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本文提出零样本概念命名协议,利用中等规模多模态大模型对局部区域进行概念标注,无需训练即可实现62%-88%的物体级精确匹配,为低成本可解释AI提供新思路。
Comments 6 pages, 2 figures, 4 tables. Accepted at the 2026 IEEE International Conference on Artificial Intelligence (CAI), 8-10 May 2026, Granada, Spain. Code: https://github.com/darianfgUgr/CoNa
Journal ref 2026 IEEE International Conference on Artificial Intelligence (CAI), Granada, Spain, 2026, pp. 1405-1410