arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 1199 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 1199 篇

2604.03455 2026-04-07 cs.IR cs.CL cs.LG 81%

Lightweight Query Routing for Adaptive RAG: A Baseline Study on RAGRouter-Bench

轻量级查询路由用于自适应RAG:对RAGRouter-Bench的基准研究

Prakhar Bansal, Shivangi Agarwal

专题命中 RAG评测 :RAG(title);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

AI总结 本文评估了轻量级分类器在RAGRouter-Bench上的路由效果,发现TF-IDF与SVM组合在宏平均F1值和准确率上表现最佳,同时实现token节省。

Comments 5 pages, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13227 2026-03-30 cs.IR cs.AI 81%

Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?

内部知识:RAG系统能从评估秘密中获得多少收益?

Laura Dietz, Bryan Li, Eugene Yang, Dawn Lawrie, William Walden, James Mayfield

机构 * University of New Hampshire(新罕布什尔大学) University of Pennsylvania(宾夕法尼亚大学) Human Language Technology Center of Excellence, Johns Hopkins University(约翰霍普金斯大学人类语言技术卓越中心)

专题命中 RAG评测 :RAG(title,abstract);分类 cs.IR、cs.AI

AI总结 本文通过对比实验探讨RAG系统在评估秘密泄露时的评估风险,指出盲评和方法多样性的重要性。

Comments To appear in ECIR 2026, Lecture Notes in Computer Science, Volume 16483

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12254 2026-03-13 cs.AI cs.IR 81%

Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation

移动代理-RAG:通过上下文知识增强驱动智能多代理协调以实现长周期移动自动化

Yuxiang Zhou, Jichang Li, Yanhao Zhang, Haonan Lu, Guanbin Li

专题命中 RAG评测 :RAG(title,abstract);分类 cs.IR、cs.AI

AI总结 本文提出Mobile-Agent-RAG框架,通过双层检索增强实现智能多代理协调,提升长周期移动自动化任务的完成率和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01710 2026-03-03 cs.CL cs.IR cs.LG 81%

Legal RAG Bench: an end-to-end benchmark for legal RAG

法律RAG基准:法律RAG的端到端评估基准

Abdur-Rahman Butler, Umar Butler

专题命中 RAG评测 :RAG(title,abstract);分类 cs.IR、cs.CL

AI总结 Legal RAG Bench通过评估法律RAG系统端到端性能,发现信息检索是性能的主要驱动因素,Kanon 2嵌入器显著提升了正确性和基础性。

Comments 13 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24277 2026-03-02 cs.IR cs.AI 81%

Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment

用于帮助读者评估新闻可信度的自动化评估资源的辅助RAG系统

Dake Zhang, Mark D. Smucker, Charles L. A. Clarke

机构 * University of Waterloo(滑铁卢大学)

专题命中 RAG评测 :RAG(title,abstract);分类 cs.IR、cs.AI

AI总结 本文介绍了用于评估辅助RAG系统以帮助读者评估新闻可信度的资源,包括问题生成和报告生成任务及自动化评估方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05713 2026-02-10 cs.CL cs.AI 81%

DRAGOn: Designing RAG On Periodically Updated Corpus

DRAGOn:设计一个定期更新语料库的RAG基准测试

Fedor Chernogorskii, Sergei Averkiev, Liliya Kudraleeva, Zaven Martirosian, Maria Tikhonova, Valentin Malykh, Alena Fenogenova

机构 * SberAI MBZUAI ITMO MISIS HSE University(俄罗斯高等经济大学) MWS AI IITU(俄罗斯联邦信息技术大学) YSDA

专题命中 RAG评测 :RAG(title,abstract);分类 cs.CL、cs.AI

AI总结 DRAGOn通过定期更新的语料库设计RAG基准测试,提供自动问题生成和评估流程,减少数据泄漏并促进社区参与。

Comments EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08847 2026-01-16 cs.CL cs.AI 81%

Scalable and Reliable Evaluation of AI Knowledge Retrieval Systems: RIKER and the Coherent Simulated Universe

可扩展且可靠的AI知识检索系统评估:RIKER和一致性模拟宇宙

JV Roig

机构 * Kamiwaza AI

专题命中 RAG评测 :knowledge retrieval(title);RAG(abstract);分类 cs.CL、cs.AI

AI总结 RIKER通过生成文档而非提取地面真实,提供了一种可扩展且抗污染的AI知识检索系统评估方法,揭示了上下文长度、跨文档聚合和事实基础能力等关键问题。

Comments 26 pages, 17 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05310 2025-10-08 cs.CL cs.AI 81%

RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts

Yining She, Daniel W. Peterson, Marianne Menglin Liu, Vikas Upadhyay, Mohammad Hossein Chaghazardi, Eunsuk Kang, Dan Roth

机构 * Carnegie Mellon University(卡内基梅隆大学) Oracle Cloud Infrastructure(Oracle 云基础设施) University of Pennsylvania(宾夕法尼亚大学)

专题命中 RAG评测 :RAG(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15727 2025-07-01 cs.NI cs.AI cs.CR cs.IR 81%

Retrieval Augmented Generation Based LLM Evaluation For Protocol State Machine Inference With Chain-of-Thought Reasoning

Youssef Maklad, Fares Wael, Wael Elsersy, Ali Hamdi

机构 * Faculty of Computer Science, MSA University(计算机科学学院,MSA大学)

专题命中 RAG评测 :retrieval augmented generation(title);RAG(abstract);分类 cs.IR、cs.AI

Comments Minor modifications in sections: abstract, introduction, background problem formulation, and conclusion. (Typos and Clarifications)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20128 2025-06-26 cs.CL cs.AI cs.LG 81%

CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation

Aashiq Muhamed

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 RAG评测 :RAG(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at LLM4Eval @ SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07671 2025-06-10 cs.CL cs.AI 81%

GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation

Ionut-Teodor Sorodoc, Leonardo F. R. Ribeiro, Rexhina Blloshmi, Christopher Davis, Adrià de Gispert

机构 * Amazon AGI(亚马逊人工智能研究院)

专题命中 RAG评测 :RAG(title,abstract);分类 cs.CL、cs.AI

Comments ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18024 2025-04-28 cs.CE cs.CL cs.IR 81%

SMARTFinRAG: Interactive Modularized Financial RAG Benchmark

Yiwei Zha

机构 * Khoury College of Computer Science(科里学院计算机科学学院) Northeastern University(东北大学)

专题命中 RAG评测 :RAG(title,abstract);分类 cs.IR、cs.CL

Comments For open source github repo, see https://github.com/JonathanZha47/SMARTFinRAG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16161 2025-03-21 cs.CL cs.AI 81%

Towards Lighter and Robust Evaluation for Retrieval Augmented Generation

Alex-Razvan Ispas, Charles-Elie Simon, Fabien Caspani, Vincent Guigue

专题命中 RAG评测 :retrieval augmented generation(title);RAG(abstract);分类 cs.CL、cs.AI

Comments 17 pages, 5 figures, published at 1st workshop of Quantify Uncertainty and Hallucination in Foundation Models: The Next Frontier in Reliable AI at ICLR 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07542 2024-08-15 cs.CY cs.AI cs.IR cs.LG 81%

New Curriculum, New Chance -- Retrieval Augmented Generation for Lesson Planning in Ugandan Secondary Schools. Prototype Quality Evaluation

Simon Kloker, Herbertson Bukoli, Twaha Kateete

专题命中 RAG评测 :retrieval augmented generation(title,abstract);分类 cs.IR、cs.AI

Comments Presented at Ndejje University Second Annual Research Dissemination Symposium 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00419 2026-08-04 cs.LG cs.AI cs.CL cs.IR 新提交 80%

Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

释放大语言模型的潜力:面向实时、企业级部署的蓝图

Muhammad Faizan Raza, Shuo, Yang, Satish Mahadevan Srinivasan, Joanna F. DeFranco

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 针对实时受监管场景中大语言模型的问题,提出基于模式的LLMOps架构,整合多模块并实现四项模式化贡献,优化权衡同时支持高风险领域的可审计与回滚部署。

Comments 6 pages, 1 figure. Authors' accepted version of an article published in IEEE Computer. The version of record is available at the DOI below

Journal ref Computer, vol. 59, no. 4, pp. 195-199, April 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24937 2026-07-28 cs.AI cs.CL cs.IR cs.LG 版本更新 80%

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

《智能体AI的搭车指南:从基础到系统》

Haggai Roitman

机构 * Haggai Roitman

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 本文为构建自主AI系统提供全面实践参考,覆盖从Transformer架构、训练优化到对齐、推理、检索增强生成、记忆系统、智能体设计模式及多智能体协调的完整技术栈。

Comments version 1.3

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03447 2026-07-07 cs.IR cs.AI cs.CL 新提交 80%

TRIAGE: Trustworthy Retrieval Instrumentation And Graph Evaluation

TRIAGE:可信检索工具与图评估

Axel TahmasebiMoradi, Lucas Schott, Martin Royer

机构 * IRT-SystemX(SystemX研究院)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 研究基于LLM驱动自动构建的知识图谱评估问题,引入TRIAGE框架,通过特定阶段指标对知识图谱构建与使用各阶段评估,可定位失败环节并给出改进方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13378 2026-07-07 cs.AI cs.CL cs.IR cs.LG q-bio.QM 版本更新 80%

DrugAgent: Reliable Multi-Agent Integration of Conflicting Biomedical Evidence for Drug-Target Interaction Assessment

DrugAgent:用于药物-靶点相互作用评估的冲突生物医学证据的可靠多智能体整合

Yoshitaka Inoue, Tianci Song, Xinling Wang, Rui Kuang, Tianfan Fu, Augustin Luna

机构 * Department of Computer Science and Engineering, University of Minnesota(计算机科学与工程系,明尼苏达大学) Computational Biology Branch, National Library of Medicine(国家医学图书馆计算生物学分支) Khoury College of Computer Sciences, Northeastern University(东北大学计算机科学学院) State Key Laboratory for Novel Software Technology at Nanjing University, School of Computer Science, Nanjing University(南京大学新型软件技术国家重点实验室,南京大学计算机科学学院) Developmental Therapeutics Branch, National Cancer Institute(国家癌症研究所发育治疗分支)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 研究药物-靶点相互作用评估中整合异构数据问题,核心方法是基于大语言模型的多智能体系统DrugAgent,贡献是能进行异构证据支持的DTI评估,补充独立DTI预测,还提供相关策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06234 2026-06-10 cs.RO cs.HC 版本更新 80%

RobotEQ: Transitioning from Passive Intelligence to Active Intelligence in Embodied AI

RobotEQ:从被动智能到主动智能的具身AI过渡

Kuofei Fang, Xinyi Che, Haomin Ouyang, Shufan Zhang, Xuehao Wang, Qi Liu, Liyi Liu, Chenqi Zhang, Wenxi Cai, Wenyu Dai, Jinyang Wu, Fan Zhang, Haoyu Chen, Bin He, Zheng Lian

机构 * State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University(自主智能无人系统国家重点实验室,同济大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) CMVS, University of Oulu(奥卢大学CMVS)

专题命中 RAG评测 :RAG(summary_cn,abstract)

AI总结 提出RobotEQ基准,评估模型在具身场景中理解并遵守社会规范的能力,实验表明现有模型在主动智能上仍有不足,利用RAG技术可提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06963 2026-05-11 cs.HC cs.AI cs.CL cs.IR 80%

From Surface Learning to Deep Understanding: A Grounded AI Tutoring System for Moodle

从表层学习到深度理解:面向Moodle的 grounded AI 教学系统

Anna Ostrowska, Michał Kukla, Gabriela Majstrak, Jan Opala, Sebastian Pergała, Jan Skwarek, Anna Wróblewska

机构 * Faculty of Mathematics and Information Sciences, Warsaw University of Technology(数学与信息科学学院,华沙技术大学)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 本文提出一种基于Moodle的AI教学助手,利用检索增强生成技术提供高质量教育,通过双核心设计实现学生互动式辅导和教师监督内容生成,有效减少信息误导并促进深度理解。

Comments 5 pages, accepted as demo paper at IJCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21745 2026-04-29 cs.DL 80%

AI-Augmented Bibliometric Framework: A Paradigm Shift with Agentic AI for Dynamic, Snippet-Based Research Analysis

增强型文献计量框架:基于代理AI的范式转变用于动态、片段式研究分析

Adela Bara, Simona-Vasilica Oprea

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval augmented generation(abstract);retriever(abstract)

AI总结 本文提出一个生成式多代理AI框架,通过自然语言指令实现动态代码基文献计量分析,无需专业编程技能,支持多模态全文检索、代理探索和动态指标创建,突破传统工具的限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27353 2026-07-31 cs.CL 新提交 79%

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

LayerRAG-Bench:面向智能体检索增强生成的跨层可靠性基准

Musa Shams

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);分类 cs.CL

AI总结 该研究推出LayerRAG-Bench跨层可靠性基准,含多领域任务与模型数据,发现模式标准化无法修复部分故障,验证了分层评估原则的重要性。

Comments 10 pages, 9 tables. Code and data: https://github.com/MusaShams/layerrag-bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02956 2026-04-14 cs.CL 79%

Enhancing Multilingual RAG Systems with Debiased Language Preference-Guided Query Fusion

通过去偏语言偏好引导查询融合增强多语言RAG系统

Jeonghyun Park, Byeongjeong Kim, Seojin Hwang, Hwanhee Lee

机构 * Chung-Ang University(中央大学)

专题命中 RAG评测 :RAG(title);retrieval-augmented generation(abstract);分类 cs.CL

AI总结 本文提出DeLP指标以消除评估基准中的结构偏见,发现多语言检索增强生成系统更倾向单语对齐,由此提出DELTA框架优化跨语言检索与生成,实验表明其在多种语言上优于英语枢纽和mRAG基线。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07964 2026-04-10 cs.AI cs.LG 79%

Are we still able to recognize pearls? Machine-driven peer review and the risk to creativity: An explainable RAG-XAI detection framework with markers extraction

我们仍然能识别珍珠吗?机器驱动的同行评审与创造力风险:一个可解释的RAG-XAI检测框架与标记提取

Alin-Gabriel Văduva, Simona-Vasilica Oprea, Adela Bâra

机构 * Bucharest University of Economic Studies(布加勒斯特经济研究大学)

专题命中 RAG评测 :RAG(title,abstract);分类 cs.AI

AI总结 本文提出RAG-XAI框架,通过标记提取检测自动化评审,以保留科学的透明度和创造力,实验显示其在检测性能上显著优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23562 2026-03-31 cs.LG cs.AI 79%

Synthetic Mixed Training: Scaling Parametric Knowledge Acquisition Beyond RAG

合成混合训练:在RAG之上扩展参数化知识获取

Seungju Han, Konwoo Kim, Chanwoo Park, Benjamin Newman, Suhas Kotha, Jaehun Jung, James Zou, Yejin Choi

机构 * Stanford University(斯坦福大学) MIT(麻省理工学院) University of Washington(华盛顿大学)

专题命中 RAG评测 :RAG(title,abstract);分类 cs.AI

AI总结 本文提出合成混合训练方法,结合合成问答和文档,通过互补训练信号提升模型性能,在QuaLITY基准上实现4.4%的相对提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09435 2026-03-11 cs.AI 79%

AI Act Evaluation Benchmark: An Open, Transparent, and Reproducible Evaluation Dataset for NLP and RAG Systems

AI行为评估基准:为NLP和RAG系统设计的开放、透明且可复现的评估数据集

Athanasios Davvetas, Michael Papademas, Xenia Ziouvelou, Vangelis Karkaletsis

专题命中 RAG评测 :RAG(title,abstract);分类 cs.AI

AI总结 本文提出了一种开放透明的数据集,用于评估NLP和RAG系统在欧盟AI法案下的合规性,通过生成风险等级分类等任务提升评估效率。

Comments 10 pages, 1 figure, 4 tables, 2 equations

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10577 2026-02-19 cs.IR 79%

Campaign-2-PT-RAG: LLM-Guided Semantic Product Type Attribution for Scalable Campaign Ranking

Campaign-2-PT-RAG: LLM引导的语义产品类型归因用于可扩展的活动排名

Yiming Che, Mansi Ranjit Mane, Keerthi Gopalakrishnan, Parisa Kaghazgaran, Murali Mohana Krishna Dandu, Archana Venkatachalapathy, Sinduja Subramaniam, Yokila Arora, Evren Korpeoglu, Sushant Kumar, Kannan Achan

专题命中 RAG评测 :RAG(title,abstract);分类 cs.IR

AI总结 Campaign-2-PT-RAG通过LLM解析活动内容并推断产品类型,生成高质量标签以提升电子商务活动排名优化

Comments fix typo and author names

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14460 2026-01-22 cs.IR 79%

Trust Me on This: A User Study of Trustworthiness for RAG Responses

相信我:对RAG响应可信度的用户研究

Weronika Łajewska, Krisztian Balog

专题命中 RAG评测 :RAG(title);retrieval-augmented generation(abstract);分类 cs.IR

AI总结 本研究通过用户实验探讨了不同解释类型如何影响用户对RAG响应的信任度,发现信任受响应清晰度和用户知识影响,而非单纯客观质量。

Comments This is the author's version of the work. The definitive version is published in: Proceedings of the 48th European Conference on Information Retrieval (ECIR '26), March 29-April 2, 2026, Delft, The Netherlands

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13671 2025-12-02 cs.CL 79%

HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings

HEALTH-PARIKSHA: 评估医疗聊天机器人在真实多语言环境中的RAG模型

Varun Gumma, Ananditha Raghunath, Mohit Jain, Sunayana Sitaram

机构 * Microsoft(微软)

专题命中 RAG评测 :RAG(title);retrieval augmented generation(abstract);分类 cs.CL

AI总结 HEALTH-PARIKSHA研究评估了24个LLMs在真实多语言环境中的表现,发现印地语模型在印地语查询中表现不一,且事实准确性低于英语查询。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10584 2025-11-05 cs.SE cs.AI 79%

ARPaCCino: An Agentic-RAG for Policy as Code Compliance

Francesco Romeo, Luigi Arena, Francesco Blefari, Francesco Aurelio Pironti, Matteo Lupinacci, Angelo Furfaro

机构 * University of Calabria(卡塔尼亚大学) IMT School for Advanced Studies Lucca(卢塞恩高级研究学院)

专题命中 RAG评测 :RAG(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏