arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

2026-03-24 至 2026-03-24 共收录 7 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 7 篇

2603.21359 2026-03-24 cs.CL cs.AI cs.CY 84%

Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF

印地语方言偏见评估:整合基于RAG的翻译和人类增强的RLAIF的多阶段框架

K. M. Jubair Sami, Dipto Sumit, Ariyan Hossain, Farig Sadeque

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一个多阶段框架,通过基于RAG的翻译和人类增强的RLAIF评估九种印地语方言的偏见。研究发现方言差异导致性能下降,模型规模增加不总是缓解偏见。

Comments 12 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20882 2026-03-24 cs.IR cs.AI cs.CL cs.LG 78%

RubricRAG: Towards Interpretable and Reliable LLM Evaluation via Domain Knowledge Retrieval for Rubric Generation

RubricRAG: 通过领域知识检索实现可解释且可靠的LLM评估

Kaustubh D. Dhole, Eugene Agichtein

机构 * Department of Computer Science(计算机科学系) Emory University(埃默里大学)

专题命中 RAG评测 :knowledge retrieval(title);分类 cs.IR、cs.CL、cs.AI

AI总结 本文提出RubricRAG,通过领域知识检索生成可解释的评估 rubric,提升LLM评估的透明性和有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02868 2026-03-24 cs.AI 70%

PrecLLM: A Privacy-Preserving Framework for Efficient Clinical Annotation Extraction from Unstructured EHRs using Small-Scale LLMs

PrecLLM: 一种用于从非结构化电子健康记录中高效提取临床注释的隐私保护框架,使用小型语言模型

Yixiang Qu, Yifan Dai, Shilin Yu, Pradham Tanikella, Malvika Pillai, Walter Chen, Jialiu Xie, Yishan Ren, Duan Wang, Yikai Wang, Sid Sheth, Guanting Chen, Yufeng Liu, Travis Schrank, Trevor Hackman, Didong Li, Di Wu

机构 * Department of Biostatistics, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校生物统计学系) Department of Genetics, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校遗传学系) Curriculum for Bioinformatics and Computational Biology, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校生物信息学与计算生物学课程) Carolina Health Informatics Program, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校健康信息学计划) Department of Statistics and Operations Research, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校统计学与运筹学系) Department of Otolaryngology/Head and Neck Surgery, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校耳鼻喉科及头颈外科系) Department of Statistics, University of Michigan(密歇根大学统计学系) Department of Biomedical Sciences, Adams School of Dentistry, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校阿德姆牙科学院生物医学科学系) Computational Medicine Program, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校计算医学计划) Lineberger Comprehensive Cancer Center, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校林伯格综合癌症中心)

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract);分类 cs.AI

AI总结 本文提出PrecLLM框架,利用小型语言模型高效处理非结构化电子健康记录,通过正则表达式和RAG技术提升隐私保护下的临床注释提取性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22184 2026-03-24 cs.LG quant-ph 67%

Revisiting Quantum Code Generation: Where Should Domain Knowledge Live?

重新审视量子代码生成:领域知识应该在哪里居住?

Oscar Novo, Oscar Bastidas-Jossa, Alberto Calvo, Antonio Peris, Carlos Kuchkovsky

机构 * Quantum Computing Research, QCentroid(量子计算研究,QCentroid)

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract)

AI总结 研究比较了Qiskit代码生成中不同专精策略,发现通用LLM在零样本和检索增强设置下表现更优,结合执行反馈代理可提升性能,表明无需领域微调即可实现更灵活的量子软件开发。

Comments Submitted to Quantum Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21329 2026-03-24 cs.IR cs.AI 62%

COINBench: Moving Beyond Individual Perspectives to Collective Intent Understanding

COINBench: 超越个体视角到集体意图理解

Xiaozhe Li, Tianyi Lyu, Siyi Yang, Yizhao Yang, Yuxi Gong, Jinxuan Huang, Ligao Zhang, Zhuoyi Huang, Qingwen Liu

机构 * Tongji University(同济大学) Stanford University(斯坦福大学) CurrentsAI Research(CurrentsAI研究院)

专题命中 RAG评测 :RAG(abstract);分类 cs.IR、cs.AI

AI总结 COINBench通过动态实时更新的基准测试,评估大语言模型在消费者领域集体意图理解能力,揭示当前模型在复杂意图合成分析深度上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12739 2026-03-24 cs.CL 57%

Edu-Values: Towards Evaluating the Chinese Education Values of Large Language Models

Edu-Values:面向评估大型语言模型的中国教育价值观

Peiyi Zhang, Yazhou Zhang, Bo Wang, Lu Rong, Prayag Tiwari, Jing Qin

专题命中 RAG评测 :RAG(abstract);分类 cs.CL

AI总结 本文提出Edu-Values基准,评估大型语言模型的中国教育价值观,包含七个核心价值,通过1418道题测试,发现中文LLM在教育文化差异下表现更优,且利用Edu-Values构建外部知识库可提升对齐效果。

Comments The authors are withdrawing this paper to make substantial revisions and improvements before future submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11721 2026-03-24 cs.AI 57%

When OpenClaw Meets Hospital: Toward an Agentic Operating System for Dynamic Clinical Workflows

当OpenClaw遇见医院:迈向动态临床工作流的智能操作系统

Wenxian Yang, Hanzheng Qiu, Bangqun Zhang, Chengquan Li, Zhiyong Huang, Xiaobin Feng, Rongshan Yu, Jiahong Dong

机构 * Independent Researcher(独立研究者) National Institute for Data Science in Health and Medicine, Xiamen University, Xiamen, China(健康医学数据科学国家研究院,厦门大学,厦门,中国) Yue’erwan Internet Hospital Co., Ltd.(悦尔湾互联网医院有限公司) School of Biomedical Engineering, Tsinghua University, Beijing, China(生物医学工程学院,清华大学,北京,中国) National University of Singapore, Singapore(新加坡国立大学) Yttrium-90 Precision Interventional Radiotherapy Center of Liver Cancer, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua University, Beijing, China(肝癌钇-90精准介入放射治疗中心,北京清华大学昌平医院,临床医学学院,清华大学,北京,中国) Hepatobiliary and Pancreatic Centre, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua University, Beijing, China(肝胆胰中心,北京清华大学昌平医院,临床医学学院,清华大学,北京,中国)

专题命中 RAG评测 :RAG(abstract);分类 cs.AI

AI总结 本文提出一种适应医院环境的LLM代理架构,通过受限执行环境、文档导向交互模型、分页索引内存架构和可组合医疗技能库,实现临床工作流的协调,保障安全、透明和可审计性。

详情

展开后加载摘要…

URL PDF HTML 收藏