arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3437 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3437 篇

2605.07655 2026-05-11 cs.CV cs.AI 81%

Towards Billion-scale Multi-modal Biometric Search

迈向十亿级多模态生物特征搜索

Arka Koner, Chetan S. Naik, Lokesh Kurre, Vivek Raghavan, Barada P. Sabut, Tanusree Deb Barma, Anoop M. Namboodiri, Anil K. Jain

机构 * Unique Identification Authority of India(印度唯一身份权威机构) IIIT Hyderabad(海得拉巴IIIT) Michigan State University(密歇根州立大学)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Bharat ABIS系统,通过多模态预处理和特征提取实现十亿级生物特征库的1:N搜索,验证了其在印度Aadhaar数据库上的高效性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05594 2026-05-08 cs.CL cs.CV cs.LG 81%

The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation

上下文的成本:减轻多模态检索增强生成中的文本偏差

Hoin Jung, Xiaoqian Wang

机构 * Elmore Family School of Electrical and Computer Engineering(埃尔莫夫家族电气与计算机工程学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文探讨了多模态大语言模型在整合检索增强生成时出现的文本偏差问题,提出BAIR方法通过恢复视觉显著性并施加位置感知惩罚来提升多模态接地和诊断可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23195 2026-04-28 cs.CV cs.AI 81%

AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval

AnalogRetriever: 为模拟电路检索学习跨模态表示

Yihan Wang, Lei Li, Yao Lai, Jing Wang, Yan Lu

机构 * Tsinghua University(清华大学) The University of Hong Kong(香港大学) University of Cambridge(剑桥大学) Nanjing University of Posts and Telecommunications(南京邮电大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出AnalogRetriever,通过构建高质量数据集和三模态检索框架,实现跨模态的模拟电路检索,实验表明其在六个方向上的Recall@1达到75.2%,显著优于现有方法。

Comments 10 pages, 7 figures. Yihan Wang and Lei Li contributed equally to this paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22885 2026-04-28 cs.CV cs.AI 81%

Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization

联邦跨模态检索与缺失模态的语义路由和适配器个性化

Hefeng Zhou, Xuan Liu, Sicheng Chen, Wutong Zhang, Wu Yan, Jiong Lou, Chentao Wu, Guangtao Xue, Wei Zhao, Jie Li

机构 * Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎-鲁克桑克鲁特研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕尔默研究实验室) Shanghai Jiao Tong University(上海交通大学) Hohai University(河海大学) Shaanghai Jiao Tong University(上海交通大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出RCSR框架,通过语义路由和适配器个性化解决联邦跨模态检索中异构数据和缺失模态的问题,提升全局检索准确性和训练稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20878 2026-04-24 cs.CL cs.CV cs.LG eess.IV 81%

AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models

AITP:通过多模态大语言模型进行交通事故责任分配

Zijin Zhou, Songan Zhang

机构 * Global Institute of Future Technology(未来技术全球研究院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出AITP模型,结合多模态链式推理和法律知识检索,解决交通事故责任分配问题,并构建DecaTARA基准测试集,验证模型在责任分配、事故检测和理解任务中的领先性能。

Journal ref CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00798 2026-04-23 cs.CV cs.AI 81%

Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering

逐步多模态搜索与推理用于知识密集型视觉问答

Changin Choi, Wonseok Lee, Jungmin Ko, Wonjong Rhee

机构 * Interdisciplinary Program in Artificial Intelligence(人工智能交叉学科项目) Department of Intelligence and Information(智能与信息系) Samsung Advanced Institute of Technology(三星先进技术研究院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出PMSR框架,通过逐步构建结构化推理轨迹提升知识获取与综合能力,实验显示在六个基准测试中提升了检索召回率和端到端回答准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19093 2026-04-22 cs.CV cs.AI 81%

Multi-modal Test-time Adaptation via Adaptive Probabilistic Gaussian Calibration

多模态测试时适应 via 自适应概率高斯校准

Jinglin Xu, Yi Li, Chuxiong Sun, Xiao Xu, Jiangmeng Li, Fanjiang Xu

机构 * Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学) National Defense University(国防大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出自适应概率高斯校准方法,解决多模态测试时适应中类别条件分布建模不足问题,通过引入自适应对比不对称修正技术,提升预测准确性和决策边界可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12735 2026-04-22 cs.CV cs.CL 81%

VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph

VimRAG:通过多模态记忆图导航大规模视觉上下文

Qiuchen Wang, Shihang Wang, Yu Zeng, Qiang Zhang, Fanrui Zhang, Zhuoning Guo, Bosi Zhang, Wenxuan Huang, Lin Chen, Zehui Chen, Pengjun Xie, Ruixue Ding

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 VimRAG通过多模态记忆图提升检索增强生成在处理大规模视觉上下文中的能力,采用动态有向无环图结构和图调制视觉记忆编码机制,实现高效的信息检索与推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16313 2026-04-21 cs.IR cs.AI cs.CL 81%

MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering

MARA:一种多模态自适应检索增强框架用于文档问答

Hui Wu, Haoquan Zhai, Yuchen Li, Hengyi Cai, Peirong Zhang, Yidan Zhang, Lei Wang, Chunle Wang, Yingyan Hou, Shuaiqiang Wang, Dawei Yin

机构 * Key Laboratory of Target Cognition and Application Technology (TCAT), AIRI, CAS(目标认知与应用技术重点实验室(TCAT),空气动力研究所,中国科学院) School of Electronic, Electrical and Communication Engineering, UCAS(电子、电气与通信工程学院,中国科学院大学) Aerospace Information Research Institute, Chinese Academy of Sciences(航天信息研究所,中国科学院) Baidu Inc.(百度公司)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 MARA框架通过引入查询自适应机制提升多模态文档检索与生成的精度和质量,实验表明其在多模态问答基准上优于现有最先进方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07415 2026-04-21 cs.CV cs.CL cs.LG 81%

RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction

RA-RRG:结合关键短语提取的多模态检索增强放射学报告生成

Jonggwon Park, Byungmu Yoon, Soobum Kim, Kyoyun Choi

机构 * DEEPNOID Inc.(DEEPNOID公司) Department of Artificial Intelligence and Data Science, Sejong University(Sejong大学人工智能与数据科学系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 RA-RRG通过结合多模态检索与大语言模型生成放射学报告,减少幻觉和计算需求,实验显示在CheXbert指标上优于现有模型。

Comments ACL 2026, Findings of the Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15210 2026-04-17 cs.AI cs.CL 81%

Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

学习像卡通标题作家一样思考:用于多模态幽默理解的不一致-解决监督

Hatice Merve Vural, Doga Kukul, Ege Erdem Ozlu, Demir Ekin Arikan, Bob Mankoff, Erkut Erdem, Aykut Erdem

机构 * Koç University(科克大学) KUIS AI Center(KUIS人工智能中心) Air Mail and Cartoon Collections(Air Mail和卡通收藏) Hacettepe University(哈恰特佩大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出IRS框架,通过分解幽默理解为不一致建模、解决建模和偏好对齐,提升多模态幽默理解能力,在NYCC基准上优于现有模型,且在零样本迁移中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12352 2026-04-15 cs.AI cs.CL 81%

MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents

多文档融合:一种用于长工业文档增强RAG的分块流程

Joongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo, Heuiseok Lim

机构 * Human-inspired AI Research(人机协同人工智能研究所) Department of Computer Science and Engineering(计算机科学与工程系) School of Software(软件学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 针对长工业文档结构复杂的问题,提出MultiDocFusion分块流程,结合视觉解析、OCR提取、层次结构重建和DFS分组,提升RAG检索精度和问答质量。

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18095 2026-04-08 cs.IR cs.CL cs.CV 81%

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction

MetaEmbed:通过灵活的后期交互实现测试时多模态检索的扩展

Zilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen, Xintao Chen, Vicente Ordonez, Vijai Mohan

机构 * Meta Superintelligence Labs(Meta超级智能实验室) Rice University(莱斯大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 MetaEmbed通过灵活的后期交互机制,在测试时实现多模态检索的扩展,通过多向量检索训练提升信息组织能力,实现在大规模模型下的高效检索性能。

Comments ICLR 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14408 2026-03-24 cs.CV cs.AI 81%

Feature Recalibration Based Olfactory-Visual Multimodal Model for Enhanced Rice Deterioration Detection

基于特征重校准的嗅觉-视觉多模态模型用于增强稻米劣变检测

Rongqiang Zhao, Hengrui Hu, Yijing Wang, Mingchun Sun, Jie Liu

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) National Key Laboratory of Smart Farm Technologies and Systems(国家智能农业技术与系统重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出基于特征重校准的嗅觉-视觉多模态模型,通过改进的特征嵌入构造器和重校准注意力网络,提升稻米劣变检测的准确性和效率,相比传统方法提升11.51%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16289 2026-03-19 cs.CV cs.AI 81%

VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents

VisBrowse-Bench:多模态浏览代理的视觉原生搜索基准测试

Zhengbo Zhang, Jinbo Su, Zhaowen Zhou, Changtao Miao, Yuhan Hong, Qimeng Wu, Yumeng Liu, Feier Wu, Yihe Tian, Yuhao Liang, Zitong Shan, Wanke Xia, Yi-Fan Zhang, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan

机构 * CASIA Ant Digital Technologies Ant Group(蚂蚁集团数字技术部) RUC(清华大学) FZU(福建师范大学) THU(清华大学) USTB(北京科技大学) PKU(北京大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出VisBrowse-Bench,用于评估多模态浏览代理的视觉推理能力,通过文本-图像检索和联合推理进行多模态证据交叉验证,验证了不同模型在搜索过程中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08202 2026-03-10 cs.CV cs.AI 81%

MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data

MM-TS: 多模态温度和边距调度用于长尾数据的对比学习

Siarhei Sheludzko, Dhimitrios Duka, Bernt Schiele, Hilde Kuehne, Anna Kukleva

机构 * University of Bonn(波恩大学) MPI for Informatics, SIC(信息研究所) Tuebingen AI Center/University of Tuebingen(图宾根人工智能中心/图宾根大学) MIT-IBM Watson AI Lab(麻省理工-IBM Watson人工智能实验室)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 MM-TS通过动态温度和边距调度提升多模态对比学习在长尾数据上的性能,实现InfoNCE和最大边距方法的统一。

Comments 18 pages, 11 figures. Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04772 2026-03-06 cs.CL cs.AI 81%

TSEmbed: Unlocking Task Scaling in Universal Multimodal Embeddings

TSEmbed: 解锁通用多模态嵌入中的任务扩展

Yebo Wu, Feng Liu, Ziwei Xie, Zhiyuan Liu, Changwang Zhang, Jun Wang, Li Li

机构 * State Key Laboratory of IOTSC, University of Macau(物联网科学与技术国家重点实验室,澳门大学) OPPO Research Institute(OPPO研究院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 TSEmbed通过融合MoE与LoRA,结合专家感知负采样策略,提升通用多模态嵌入的任务扩展能力,并在基准和工业数据集上取得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03637 2026-03-05 cs.CV cs.AI cs.CR 81%

Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions

基于图像的提示注入:通过视觉嵌入的对抗性指令劫持多模态大语言模型

Neha Nagaraja, Lan Zhang, Zhilong Wang, Bo Zhang, Pawan Patil

机构 * Northern Arizona University(北亚利桑那大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 研究提出基于图像的提示注入攻击方法,通过视觉嵌入对抗性指令,成功操控多模态大语言模型输出,揭示了该攻击的有效性和潜在威胁。

Comments 7 pages, published in 2025 3rd International Conference on Foundation and Large Language Models (FLLM), Vienna, Austria

Journal ref 2025 3rd International Conference on Foundation and Large Language Models (FLLM), Vienna, Austria, 2025, pp. 916-922

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23369 2026-03-02 cs.IR cs.AI cs.CL 81%

Reason to Contrast: A Cascaded Multimodal Retrieval Framework

对比之理:一种级联多模态检索框架

Xuanming Cui, Hong-You Chen, Hao Yu, Hao Yuan, Zihao Wang, Shlok Kumar Mishra, Hanchao Yu, Yonghuan Yang, Jun Xiao, Ser-Nam Lim, Jianpeng Cheng, Qi Guo, Xiangjun Fan

机构 * AI at Meta(Meta人工智能部门) University of Central Florida(中央佛罗里达大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出TTE-v2框架,通过引入基于额外输入标记预算的推理驱动性能扩展,提升多模态检索的测试性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01390 2026-02-16 cs.CV cs.AI cs.LG 81%

Multimodal Doctor-in-the-Loop: A Clinically-Guided Explainable Framework for Predicting Pathological Response in Non-Small Cell Lung Cancer

多模态医生在环:一种临床指导的可解释框架,用于预测非小细胞肺癌的病理反应

Alice Natalina Caragliano, Claudia Tacconi, Carlo Greco, Lorenzo Nibid, Edy Ippolito, Michele Fiore, Giuseppe Perrone, Sara Ramella, Paolo Soda, Valerio Guarrasi

机构 * Research Unit of Computer Systems and Bioinformatics, Department of Engineering(计算机系统与生物信息学研究单位,工程系) Research Unit of Radiation Oncology, Department of Medicine and Surgery(放射肿瘤学研究单位,医学与外科系) Research Unit of Anatomical Pathology, Department of Medicine and Surgery(病理学研究单位,医学与外科系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究提出多模态医生在环框架,结合影像与临床数据,通过嵌入医生知识提升肺癌病理反应预测的准确性和可解释性。

Comments arXiv admin note: substantial text overlap with arXiv:2502.17503

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11737 2026-02-13 cs.CV cs.CL 81%

Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding

关注关键要素:通过对象对齐的视觉对比解码缓解多模态大语言模型中的对象幻觉

Boqi Chen, Xudong Liu, Jianing Qiu

机构 * ETH Zurich(苏黎世联邦理工学院) Amazon(亚马逊) MBZUAI(穆罕默德·本·拉什德智能科学技术研究院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出了一种通过对象对齐的视觉对比解码缓解多模态大语言模型中对象幻觉的方法,具有计算开销低且兼容性强的特点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08342 2026-02-10 cs.CV cs.AI 81%

UrbanGraphEmbeddings: Learning and Evaluating Spatially Grounded Multimodal Embeddings for Urban Science

UrbanGraphEmbeddings: 学习和评估空间导向的多模态嵌入用于城市科学

Jie Zhang, Xingtong Yu, Yuan Fang, Rudi Stouffs, Zdravko Trivic

机构 * National University of Singapore Department of Architecture Singapore The Chinese University of Hong Kong Dept of Systems Eng. \& Eng. Mgmt. China Singapore Management University School of Computing \& Info. Systems Singapore National University of Singapore Department of Architecture The Chinese University of Hong Kong Singapore Management University

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出UGE框架,通过空间导向的多模态嵌入提升城市理解任务性能,实验显示在图像检索和地理位置排名上取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04136 2026-02-04 cs.CV cs.AI 81%

UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval

UniFGVC: 一种通用无训练少样本细粒度视觉分类方法通过属性感知多模态检索

Hongyu Guo, Xiangzhao Hao, Jiarui Guo, Haiyun Guo, Jinqiao Wang, Tat-Seng Chua

机构 * School of Traffic and Transportation, Beijing Jiaotong University(交通与运输学院,北京交通大学) Foundation Modal Research Center, Institute of Automation, Chinese Academy of Sciences(基础模态研究中心,自动化研究所) Queen Mary School Hainan, Beijing University of Posts and Telecommunications(海南女王学院,北京邮电大学) Department of Computer Science, NUS School of Computing(计算机科学系,国立大学计算机学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 UniFGVC通过属性感知多模态检索方法,实现无训练少样本细粒度视觉分类,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02004 2026-02-03 cs.CV cs.AI 81%

ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning

ClueTracer: 问题到视觉线索追踪用于无训练 hallucination 抑制在多模态推理

Gongli Xi, Kun Wang, Zeming Gao, Huahui Yi, Haolang Lu, Ye Tian, Wendong Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学) West China Biomedical Big Data Center(西京生物大数据中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 ClueTracer通过问题到视觉线索追踪,无训练抑制多模态推理中的幻觉,提升推理和非推理任务性能。

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01561 2026-02-03 cs.CV cs.AI 81%

Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Odd

多模态UNcommonsense:从奇特到普通和从普通到奇特

Yejin Son, Saejin Kim, Dongjun Min, Younjae Yu

机构 * Yonsei University(延世大学) Seoul National University(首尔国立大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 多模态UNcommonsense通过R-ICL框架提升模型在非典型场景下的推理能力,实现从奇特到普通和从普通到奇特的转换。

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19060 2026-01-28 cs.CV cs.AI 81%

Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models

基于像素的检索用于知识型大型多模态模型

Jeonghwan Kim, Renjie Tao, Sanat Sharma, Jiaqi Wang, Kai Sun, Zhaojiang Lin, Seungwhan Moon, Lambert Mathias, Anuj Kumar, Heng Ji, Xin Luna Dong

机构 * Meta Reality Labs(Meta现实实验室) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 PixSearch是一种端到端的多模态模型,通过像素级检索和统一感知与推理,提升视觉问答任务中的事实一致性与泛化能力。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02729 2026-01-26 cs.CL cs.AI cs.IR 81%

Unified Multimodal Interleaved Document Representation for Retrieval

统一多模态交错文档表示用于检索

Jaewoo Lee, Joonho Ko, Jinheon Baek, Soyeong Jeong, Sung Ju Hwang

机构 * University of North Carolina Chapel Hill(北卡罗来纳大学教堂山分校) KAIST(韩国科学技术院) DeepAuto

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出统一多模态交错文档表示方法,通过整合文本、图像和表格信息提升信息检索性能。

Comments EACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07780 2026-01-21 cs.CV cs.AI 81%

Semantic-Consistent Bidirectional Contrastive Hashing for Noisy Multi-Label Cross-Modal Retrieval

语义一致的双向对比哈希用于噪声多标签跨模态检索

Likang Peng, Chao Su, Wenyuan Wu, Yuan Sun, Dezhong Peng, Xi Peng, Xu Wang

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出语义一致的双向对比哈希方法,通过跨模态语义一致性分类和双向软对比哈希模块,有效应对多标签数据中的噪声问题,提升跨模态检索的鲁棒性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05600 2026-01-12 cs.CV cs.CL cs.LG 81%

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

SceneAlign: 在复杂视觉场景中将多模态推理对齐到场景图

Chuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu, Chengkai Huang, Lina Yao, Julian McAuley, Jingbo Shang

机构 * University of California, San Diego(加州大学圣地亚哥分校) University of Toronto(多伦多大学) University of New South Wales(新南威尔士大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 SceneAlign通过利用场景图进行结构干预,提升多模态推理在复杂视觉场景中的准确性和忠实性。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04571 2026-01-09 cs.AI cs.MM 81%

Enhancing Multimodal Retrieval via Complementary Information Extraction and Alignment

通过互补信息提取与对齐增强多模态检索

Delong Zeng, Yuexiang Xie, Yaliang Li, Ying Shen

机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(1 智能系统工程学院,中山大学) Alibaba Group(2 阿里巴巴集团) Guangdong Provincial Key Laboratory of Fire Science and Intelligent Emergency Technology(3 广东省火灾科学与智能应急技术重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI、cs.MM

AI总结 CIEA通过互补信息提取与对齐提升多模态检索效果,实现对图像和文本统一潜在空间的建模,并在多个基准上取得显著优势。

Comments Accepted by ACL'2025

详情

展开后加载摘要…

URL PDF HTML 收藏