arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

2026-04-15 至 2026-04-15 共收录 3 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 3 篇

2604.12352 2026-04-15 cs.AI cs.CL 90%

MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents

多文档融合:一种用于长工业文档增强RAG的分块流程

Joongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo, Heuiseok Lim

机构 * Human-inspired AI Research(人机协同人工智能研究所) Department of Computer Science and Engineering(计算机科学与工程系) School of Software(软件学院)

专题命中 多模态RAG :RAG(title,title_cn);分类 cs.CL、cs.AI

AI总结 针对长工业文档结构复杂的问题,提出MultiDocFusion分块流程,结合视觉解析、OCR提取、层次结构重建和DFS分组,提升RAG检索精度和问答质量。

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05818 2026-04-15 cs.CV cs.CL cs.IR 82%

WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering

WikiSeeker: 重新思考视觉语言模型在基于知识的视觉问答中的作用

Yingjian Zhu, Xinming Wang, Kun Ding, Ying Wang, Bin Fan, Shiming Xiang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室 (MAIS))

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.IR、cs.CL

AI总结 本文提出WikiSeeker框架,通过引入多模态检索器和重新定义视觉语言模型的角色,提升多模态检索性能和答案质量,实现在EVQA、InfoSeek和M2KR数据集上的最优表现。

Comments Accepted by ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13021 2026-04-15 cs.CV cs.AI 77%

Representation geometry shapes task performance in vision-language modeling for CT enterography

表征几何形状任务性能在CT结肠镜视觉-语言建模中的作用

Cristian Minoccheri, Emily Wittrup, Kayvan Najarian, Ryan Stidham

机构 * Gilbert S. Omenn Department of Computational Medicine and Bioinformatics(Gilbert S. Omenn 计算医学与生物信息学部门) Department of Gastroenterology(消化内科部门) Department of Emergency Medicine(急诊医学部门) Department of Electrical Engineering and Computer Science(电气工程与计算机科学部门)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文研究了CT结肠镜视觉-语言迁移学习,发现均值池化在疾病分类中表现更优,而注意力池化在跨模态检索中更有效,同时指出单片组织对比度比空间覆盖范围更重要,为构建体积医学影像视觉-语言系统提供了基础。

详情

展开后加载摘要…

URL PDF HTML 收藏