D$^2$TV: Dual Knowledge Distillation and Target-oriented Vision Modeling for Many-to-Many Multimodal Summarization
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments EMNLP 2023 Findings
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments EMNLP 2023 Findings
专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments arXiv admin note: text overlap with arXiv:2208.00877 by other authors
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)
Comments 31 pages, 15 figures, and 15 tables
专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract)
Comments This paper has been published as a full paper at WWW 2023
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Published in ICML 2023. Project page: https://jykoh.com/fromage
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments ICCCI 2023
专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Journal ref Advances in Information Retrieval: 45th European Conference on Information Retrieval, ECIR 2023, Dublin, Ireland, Proceedings, Part II
专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted by CVPR 2023
专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract)
Comments Dataset available at https://github.com/hsiehjackson/Mr.Right
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments 11 pages, 4 figures
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted to 26th International Conference on Pattern Recognition (ICPR) 2022
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS
Comments Accepted by ICASSP 2022
专题命中 跨模态检索 :cross-modal(title);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract)
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Submitted for review at CIKM 2019
专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract)
专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract)
专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract)
专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract)
专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM
专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments To appear in AAAI 2018
专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM
Comments 13 pages, submitted to IEEE Transactions on Image Processing
专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract)
Comments 14 pages; Accepted by Computer Vision and Image Understanding
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)
超越上下文的规模:文档理解的多模态检索增强生成综述
机构 * MBZUAI ; Alibaba Group(阿里巴巴集团) ; Tsinghua University(清华大学) ; Wuhan University(武汉大学) ; Nanyang Technological University(南洋理工大学) ; University of Melbourne(墨尔本大学)
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL
AI总结 本文综述了多模态检索增强生成在文档理解中的应用,提出基于领域、检索模态和粒度的分类体系,总结了关键数据集、基准测试和行业应用,指出了效率、细粒度表示和鲁棒性等挑战。
Comments Accepted by ACL2026 Main Conference; Project is available at https://github.com/SensenGao/Multimodal-RAG-Survey-For-Document
FUSE : 多模态搜索与推荐中子代理证据的故障感知使用
机构 * Adobe Inc.(Adobe公司)
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI
AI总结 FUSE通过上下文压缩策略提升多模态搜索与推荐性能,实现93.3%的意图准确率和99.4%的召回率。
Comments ICDM MMSR 2025: Workshop on Multimodal Search and Recommendations
迈向食品图像到食谱检索的无偏跨模态表示学习
专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.MM
AI总结 本文通过因果理论解决食品图像到食谱检索中的偏见问题,提出因果干预方法和多标签分类器,提升检索性能。
Comments Code link: https://github.com/GZWQ/Towards-Unbiased-Cross-Modal-Representation-Learning-for-Food-Image-to-Recipe-Retrieval
机构 * Cloudglue
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Accepted to ICCV 2025 Multimodal Representation and Retrieval Workshop
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Visual Question Answering, Rank VQA, Faster R-CNN, BERT, Multimodal Fusion, Ranking Learning, Hybrid Training Strategy
专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI
Comments 5 pages, 3 figures
Journal ref [1] Chen X , Zhang R , Zhan Y . Graph Pattern Loss Based Diversified Attention Network For Cross-Modal Retrieval[C]// 2020 IEEE International Conference on Image Processing (ICIP). IEEE, 2020