arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

2026-06-11 至 2026-06-11 共收录 5 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 5 篇

2606.11199 2026-06-11 cs.CL cs.AI cs.IR cs.LG 新提交 91%

NightFeats @ MMU-RAGent NeurIPS 2025: A Context-Optimized Multi-Agent RAG System for the Text-to-Text Track

NightFeats @ MMU-RAGent NeurIPS 2025: 面向文本到文本轨道的上下文优化多智能体RAG系统

Quentin Fever, Naziha Aslam

机构 * NightFeats

专题命中 RAG评测 :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 提出一种结构化多智能体RAG系统NightFeats,通过检索、策展和组合三阶段分解知识合成,引入时序语义重排序、矛盾协调和引用保留架构,在MMU-RAGent竞赛中超越商业基线。

Comments 5 pages, 1 figure, 1 table. NeurIPS 2025 Competition Track (MMU-RAGent). System developed October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11257 2026-06-11 cs.CL cs.LG cs.PF 新提交 91%

Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X Elite

移动NPU上的能效型设备端RAG:Snapdragon X Elite系统设计与基准测试

Zhiyuan Cheng, Longying Lai

机构 * Qualcomm(高通) Snapdragon X Elite(骁龙X Elite) Dell XPS 13 laptop(戴尔XPS 13笔记本电脑) Qualcomm Hexagon NPU(高通Hexagon NPU) Adreno X1-85

专题命中 RAG评测 :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.CL

AI总结 本文首次在Snapdragon X Elite的Hexagon NPU上实现端到端RAG流水线,通过对比CPU和GPU,NPU在嵌入吞吐量、系统能耗和查询延迟上分别提升9.1倍、降低12.3倍和4.0倍,且答案质量相当。

Comments 9 pages, 2 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03792 2026-06-11 cs.CL 版本更新 77%

VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation

VietMed-MCQ:面向越南传统医学评估的一致性过滤数据合成框架

Huynh Trung Kiet, Dao Sy Duy Minh, Nguyen Dinh Ha Duong, Le Hoang Minh Huy, Long Nguyen, Dien Dinh

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL

AI总结 提出基于检索增强生成和一致性过滤的VietMed-MCQ数据集,含3190道多选题,经专家验证准确率94.2%,基准测试显示通用模型优于越南语模型。

Comments The authors have withdrawn this article because the current version is still undergoing substantial revision. Several components of the data synthesis framework, consistency-filtering procedure, evaluation protocol, and experimental analysis are being refined and expanded. As a result, the current manuscript should not be considered a complete or final representation of the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22025 2026-06-11 cs.CL cs.AI cs.IR cs.SE 版本更新 75%

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

当通用提示改进有害:LLM应用的评估驱动迭代

Daniel Commey

机构 * Daniel Commey

专题命中 RAG评测 :RAG(abstract,abstract_cn);分类 cs.IR、cs.CL、cs.AI

AI总结 提出最小可行评估套件(MVES),通过结构化评估框架和本地复现实验,发现通用提示添加并非单调改进,强调评估驱动的提示迭代。

Comments Technical report. 42 pages, 3 figures. Code, test suites, and result logs: https://github.com/dcommey/llm-eval-benchmarking

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11712 2026-06-11 cs.CL cs.AI cs.LG 新提交 73%

Substrate Asymmetry in User-Side Memory: A Diagnostic Framework

用户侧记忆中的子模块不对称性:一个诊断框架

Youwang Deng

机构 * EpistemicaLab — Independent Research(EpistemicaLab — 独立研究)

专题命中 RAG评测 :RAG(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 提出一个诊断框架,将LLM用户侧记忆分解为行为一致性、事实存在和事实缺失三个正交子模块,发现参数记忆与检索记忆在不同子模块上存在不对称性,且RLHF调优加剧了这种不对称性。

Comments Preprint. Code: https://github.com/EpistemicaLab/substrate-asymmetry-memory

详情

展开后加载摘要…

URL PDF HTML 收藏