Learning to Rematch Mismatched Pairs for Robust Cross-Modal Retrieval
专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments CVPR 2024
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments CVPR 2024
专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted by AAAI 2024
Journal ref AAAI 2024
专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM
Comments Findings of EMNLP 2023
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments EMNLP 2023 Findings
专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM
Comments This is the ArXiv version of our paper accepted by NLPCC 2023. The code will be released soon
专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted full paper at ACL 2023; 15 pages, 7 figures
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM
Comments accepted at AAAI23
专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to EMNLP 2022 main conference
专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.MM
专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments 7 pages, 2 figures
专题命中 跨模态检索 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.MM
机构 * Computer Engineering Department, Sharif University of Technology(沙斐大学计算机工程系) ; College of Interdisciplinary Science and Technology, University of Tehran(德黑兰大学跨学科科学与技术学院) ; Computer Engineering Department, K.N. Toosi University of Technology(技术学院计算机工程系) ; Qatar Computing Research Institute(卡塔尔计算研究院)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI
Comments GitHub repository: https://github.com/llm-lab-org/Multimodal-RAG-Survey
专题命中 跨模态检索 :multimodal(title);cross-modal(abstract,comments);image-text(abstract);分类 cs.CV、cs.CL
Comments Accepted at ICCV 2019 Workshop on Cross-Modal Learning in Real World
机构 * Energy and Natural Resources Security Group(能源与自然资源安全组) ; Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)
专题命中 跨模态检索 :multimodal(title);multimodal foundation model(title);分类 cs.AI
Comments Accepted at the Neurips 2025 AI4Science Workshop
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)
Comments CVPR 2024 Workshop on What is Next in Multimodal Foundation Models
专题命中 跨模态检索 :cross-modal(title);image-text(title);分类 cs.CV
Comments 9 pages; Accepted by CVPR2020
VERDICT:基于分歧感知共识的多模态推理无训练逐步验证方法
机构 * Indian Institute of Technology, Hyderabad(印度理工学院海得拉巴分校) ; Microsoft Research(微软研究院)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 该研究针对多模态大语言模型推理链易出错的问题,提出无训练的VERDICT方法,利用跨模态分歧的闭式解计算共识评分,在六个基准上提升最高达5.95%,性能接近需大量监督的特定领域评论家。
Comments European Conference on Computer Vision 2026
Conan-embedding-v3: 融合模态特定模型实现全模态嵌入
机构 * Tencent(腾讯)
专题命中 跨模态检索 :omni-modal(title,abstract);multi-modal(abstract);分类 cs.AI、cs.MM
AI总结 提出解耦-融合-恢复框架,通过独立训练模态专家并融合任务向量,再使用投影器恢复和平衡多模态重演解决投影器漂移问题,实现单一骨干网络支持文本、图像、视频、文档和音频检索。
MultiMem: 多模态对比学习中的记忆化测量与缓解
机构 * CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)
专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 提出首个多模态对比学习记忆化度量MultiMem,发现跨模态语义错位是主要驱动因素,文本模态影响最大,并通过跨模态增强有效降低记忆化并提升性能。
Comments Accepted at The 19th European Conference on Computer Vision (ECCV), 2026
多模态大语言模型中通过CoRe头的功能稀疏性机制洞察
机构 * Soochow University(苏州大学) ; Peking University(北京大学)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI
AI总结 通过识别和分析CoRe头,揭示多模态大语言模型在跨模态检索中功能稀疏的结构特性,并验证其必要性及加速推理的潜力。
阅读,而非思考:理解并弥合多模态大语言模型中文本变为像素时的模态差距
机构 * Johns Hopkins University(约翰霍普金斯大学) ; Amazon(亚马逊) ; New York University(纽约大学) ; Texas A&M University(德克萨斯大学)
专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract_cn);分类 cs.CV、cs.CL
AI总结 本文系统诊断多模态大语言模型在处理图像文本时的模态差距,发现其源于模型推理意愿不足而非感知失败,并提出一种轻量级自蒸馏方法有效弥合该差距。
xModel-KD:基于LiDAR的3D场景感知跨模态知识蒸馏
机构 * Dept. of Computer Science Lakehead University Thunder Bay, Canada ; School of Computer Science Engg. \& Info. Systems Vellore Institute of Technology Tamil Nadu, India ; Dept. of Software Engg. Lakehead University Thunder Bay, Canada
专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
AI总结 提出跨模态知识蒸馏框架xModel-KD,通过对比学习对齐2D图像纹理与3D点云几何特征,在无额外标注下提升LiDAR点云分割性能。
Comments 3 figures, and 5 tables
基于多模态知识图谱和可靠性引导精化的病例感知医学图像分类
机构 * University of Science and Technology of China(科学技术大学)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 提出一种基于多模态知识图谱的病例感知推理框架,通过构建结构化诊断记忆、自适应检索相似病例、知识传播与注入机制以及置信度校准的决策精化方案,提升医学图像分类的性能和可解释性。
对比多模态超图推理用于3D人群网格重建
机构 * Tianjin University(天津大学) ; Nanyang Technological University(南洋理工大学) ; Sichuan University(四川大学)
专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM
AI总结 本文提出对比多模态超图推理方法,结合语义、几何和姿态线索,解决人群重建中的遮挡和深度模糊问题,通过超图建模高阶人群动态,提升特征融合与跨模态正交性。
Comments ICME 2026
基于案例相似性搜索的医学影像印象 grounded 多模态检索增强草稿生成
机构 * Independent AI Researcher(独立AI研究员)
专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI
AI总结 本文提出一种多模态检索增强生成系统,结合对比图像-文本嵌入、基于案例的相似性检索和引用约束草稿生成,以生成具有临床依据的医学影像印象。
Comments 15 pages, 4 figures, 3 tables
MiMIC: 缓解通用多模态检索中的视觉模态崩溃并避免语义错位
机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) ; School of Artificial Intelligence(人工智能学院)
专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
AI总结 MiMIC通过融合-解码器架构和鲁棒训练方法,解决通用多模态检索中视觉模态崩溃和语义错位问题,在WebQA+和EVQA+数据集上优于早融合和晚融合基线。
多模态大语言模型的索引用于大规模图像检索
机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) ; VRG, FEE, Czech Technical University in Prague(布拉格捷克技术大学VRG学院)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL
AI总结 本文探讨了多模态大语言模型作为无训练相似性估计器在实例级图像到图像检索中的应用,通过提示配对图像并转换下一个词的概率为相似性分数,实现了零样本重排序,展示了其在大规模检索中的优势。
M$^3$KG-RAG:多跳多模态知识图谱增强的检索增强生成
机构 * Korea University(高丽大学) ; Sungkyunkwan University(成均馆大学) ; NVIDIA Research(英伟达研究院) ; Hanwha Systems(韩华系统)
专题命中 跨模态检索 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CL、cs.AI
AI总结 本文提出M$^3$KG-RAG,通过构建多跳多模态知识图谱并引入GRASP机制,提升多模态大语言模型的推理深度和回答准确性。
Comments Accepted to CVPR 2026
别让我看图片:通过视觉提示注入防止多模态大语言模型分析图片
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Duke University(杜克大学)
专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 本文提出ImageProtector,通过在图像中嵌入精心设计的微小扰动,使多模态大语言模型在分析时产生拒绝响应,同时评估了三种潜在的防御措施,发现它们在降低ImageProtector效果的同时影响模型性能。
Comments Appeared in ACL 2026 main conference
Journal ref The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)
MultiFinRAG:一种优化的多模态检索增强生成(RAG)框架,用于金融问答
机构 * S&P Global Ratings(标普全球评级)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI
AI总结 MultiFinRAG通过多模态提取和分层回退策略,提升金融问答任务中跨模态推理的准确性,比ChatGPT-4o在复杂任务上高19个百分点。
Comments Preprint Copy
Journal ref 2025 IEEE International Conference on Big Data (BigData), Macau, China