arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

2026-08-04 至 2026-08-04 共收录 6 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 6 篇

2608.00006 2026-08-04 cs.AI 新提交 91%

Enhancing LLMs with Context-Specific Knowledge for Mitigating Misinformation in SMEs: A RAG-based Modeling and Analysis

用特定上下文知识增强大语言模型以缓解中小企业的错误信息:一种基于RAG的建模与分析

Md. Samiul Islam, Iqbal H. Sarker, Chadni Islam, Ahmad Mohsin, Ahmed Ibrahim, Helge Janicke

机构 * School of Science, Edith Cowan University(伊迪斯科文大学理学院)

专题命中 RAG评测 :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 该研究针对中小企业使用LLM时面临的错误信息问题,提出VectorRAG与GraphRAG方法,在LLaMA等模型上验证RAG可提升响应质量,助力可靠决策。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04998 2026-08-04 cs.AI 89%

Mathematical Reasoning for Unmanned Aerial Vehicles: A RAG-Based Approach for Complex Arithmetic Reasoning

无人机数学推理:基于检索增强生成的方法用于复杂算术推理

Mehdi Azarafza, Mojtaba Nayyeri, Faezeh Pasandideh, Steffen Staab, Achim Rettberg

机构 * Department of Computer Science(计算机科学系) Institute For Artificial Intelligence(人工智能研究所) Hamm-Lippstadt University of Applied Sciences(哈姆-利普施塔特应用科学大学) University of Stuttgart(斯图加特大学)

专题命中 RAG评测 :RAG(title,summary_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出RAG-UAV框架,通过引入领域文献提升LLM在无人机任务中的数学推理能力,实验表明检索显著提高准确率并减少错误选择。

Comments 15 pages, 7 figures, 4 appendix subsections

Journal ref ICLR 2026 Workshop on Logical Reasoning of Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00824 2026-08-04 cs.SE 新提交 83%

Structure-Aware Semantic Chunking with Title-Chain Prefixes: A 1600-Query Evaluation and the Measurement Trap in Text-Transform Ablations

结合标题链前缀的结构感知语义分块:1600次查询评估及文本变换消融实验中的测量陷阱

Yang Yang

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 本文提出仅在分块侧的三阶段语义分块流水线,经1600次查询评估提升RAG的MRR@5,还发现分块研究中存在的测量陷阱并推荐检索时每个候选前缀评估的方向。

Comments 8 pages, 1 table. Replication package: DOI 10.5281/zenodo.21744653

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00419 2026-08-04 cs.LG cs.AI cs.CL cs.IR 新提交 80%

Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

释放大语言模型的潜力:面向实时、企业级部署的蓝图

Muhammad Faizan Raza, Shuo, Yang, Satish Mahadevan Srinivasan, Joanna F. DeFranco

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 针对实时受监管场景中大语言模型的问题,提出基于模式的LLMOps架构,整合多模块并实现四项模式化贡献,优化权衡同时支持高风险领域的可审计与回滚部署。

Comments 6 pages, 1 figure. Authors' accepted version of an article published in IEEE Computer. The version of record is available at the DOI below

Journal ref Computer, vol. 59, no. 4, pp. 195-199, April 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01783 2026-08-04 cs.SD cs.HC stat.AP 新提交 75%

Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias

GPT-4o-mini与教师平均评分在音乐分析回答自动评分中的比较验证:单次部署、可重复性及策略特异性偏差

Baicheng Lin, Lingxi Jin, Kyung-Seok Min

机构 * Sejong University(世宗大学) Ewha Womans University(梨花女子大学)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract)

AI总结 本研究以教师平均评分为基准,评估GPT-4o-mini对音乐分析回答的自动评分能力,发现Fs+CoT策略与教师评分一致性最强,不同策略评分特征不同,实际应用需针对性校准与人工监督。

Comments Proceedings of the International Meeting of the Psychometric Society: The 91st Annual Meeting, Seoul, Republic of Korea, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00009 2026-08-04 cs.CL cs.AI 新提交 62%

AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents

AgentMemBench:用于评估对话AI智能体长期记忆管理策略的系统基准

Ahmed Cherif

机构 * Sofrecom(索弗雷科姆)

专题命中 RAG评测 :dense retrieval(abstract);分类 cs.CL、cs.AI

AI总结 该研究构建了AgentMemBench基准,评估五种记忆策略及两款现有系统,发现外部键值存储(EKV)在长程对话记忆任务中表现最优,但存在内存占用成本,同时公开了全部可复现资源。

Comments 22 pages, 3 figures submitted on Neural Computing and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏