arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 1195 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 1195 篇

2605.15109 2026-05-15 cs.AI cs.IR 87%

Why Neighborhoods Matter: Traversal Context and Provenance in Agentic GraphRAG

为什么社区重要:代理图RAG中的遍历上下文与溯源

Riccardo Terrenzi, Maximilian von Zastrow, Serkan Ayvaz

机构 * Centre for Industrial Software, University of Southern Denmark, Alsion 2, 6400 Sønderborg, Denmark(丹麦南部大学工业软件中心)

专题命中 RAG评测 :RAG(title_cn,summary_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文探讨了代理图RAG中引用忠实性的轨迹层面问题,通过实验表明引用证据和遍历上下文都对答案准确性有影响。

Comments 7 pages, 2 figures, Submitted at IJCAI-ECAI 2026 Joint Workshop on GENAIK and NORA

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28358 2026-06-30 cs.IR cs.AI cs.CL 87%

How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation

LLM如何引用?检索增强生成中归因的机制解释

Ian van Dort, Maria Heuss

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract,abstract_cn);分类 cs.IR、cs.CL、cs.AI

AI总结 通过激活修补方法,发现LLM的引用机制并非单一组件,而是由注意力头和MLP层组成的分布式“归因集成”,调控这些组件可修复大部分错误引用。

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in Advances in Information Retrieval, ECIR 2026, Lecture Notes in Computer Science, vol. 16485, pp. 458-473, and is available online at https://doi.org/10.1007/978-3-032-21324-2_35

Journal ref Advances in Information Retrieval, ECIR 2026. Lecture Notes in Computer Science, vol. 16485, pp. 458-473. Springer, Cham (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22917 2025-08-01 cs.CL cs.AI cs.IR 87%

Reading Between the Timelines: RAG for Answering Diachronic Questions

Kwun Hang Lau, Ruiyuan Zhang, Weijie Shi, Xiaofang Zhou, Xiaojun Cheng

机构 * Hong Kong University of Science and Technology(香港科技大学) China Unicom (Hong Kong) Operation Ltd(中国移动(香港)运营有限公司)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.IR、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12322 2024-12-18 cs.LG cs.AI cs.CL cs.IR 87%

RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems

Ioannis Papadimitriou, Ilias Gialampoukidis, Stefanos Vrochidis, Ioannis, Kompatsiaris

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);vector search(abstract);分类 cs.IR、cs.CL、cs.AI

Comments Work In Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15297 2026-07-31 cs.CL 版本更新 86%

AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports

AfriEconQA:基于世界银行报告的非洲经济分析基准数据集

Edward Ajayi, Mustapha Alaba, David Stephen

机构 * Carnegie Mellon University Africa(卡内基梅隆大学非洲分校)

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);hybrid retrieval(abstract);分类 cs.CL

AI总结 AfriEconQA是一个基于世界银行报告的非洲经济分析基准数据集,旨在测试信息检索和RAG系统在处理复杂经济查询时的性能。

Comments Dataset Explorer: https://afrieconqa.pages.dev/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22706 2026-07-28 cs.AI cs.DL cs.IR 新提交 86%

MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation

MPR-CiteG:通过多组合检索和引用基础生成增强RAG

Hyewon Lee, Minkyung Song, Junghyun Oh, Seunghoon Han, Sungsu Lim

机构 * Chungnam National University(忠南国立大学) Data Intelligence Laboratory(数据智能实验室)

专题命中 RAG评测 :RAG(title,title_cn);retriever(abstract);分类 cs.IR、cs.AI

AI总结 该研究针对生成式AI检索低效与缺来源验证问题,提出双组件MPR-CiteG框架,用多组合检索器高效检索,引用基础生成模块确保输出事实一致且有源可溯,经实验验证其有效性与可靠性,助力构建更可信准确的大语言模型。

Comments 12 pages, 1 figure, 7 tables. The 1st International Workshop on Retrieval-Driven Generative AI & ScienceON AI Challenge 2025@CIKM

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25572 2026-07-24 cs.CL cs.AI 版本更新 86%

PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation

PennySynth:基于RAG的数据合成用于自动量子代码生成

Minghao Shao, Nouhaila Innan, Hariharan Janardhanan, Muhammad Kashif, Alberto Marchisio, Muhammad Shafique

机构 * eBRAIN Lab, Division of Engineering, New York University Abu Dhabi (NYUAD)(eBRAIN实验室,工程系,纽约大学阿布扎比分校) Center for Quantum and Topological Systems (CQTS), NYUAD Research Institute(量子与拓扑系统中心(CQTS),NYUAD研究所) Department of Computer Science and Engineering, NYU Tandon School of Engineering(计算机科学与工程系,纽约大学坦顿工程学院)

专题命中 RAG评测 :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 提出PennySynth框架,通过检索增强生成和代码感知嵌入,利用13,389个PennyLane指令-代码对数据集,在QHack竞赛中实现52%-68%的pass@5,显著提升量子代码生成的结构有效性和功能正确性。

Comments Accepted at the IEEE International Conference on Quantum Computing and Engineering (QCE), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09349 2026-07-13 cs.CL cs.AI cs.LG 新提交 86%

Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation

欺骗性基础:临床检索增强生成中的实体归因失败

Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim

机构 * Lunit(鲁尼特)

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 研究临床检索增强生成中的欺骗性基础问题,通过对13个模型测试发现DG率8%-87%,医学等微调模型高达86.7%。控制消融确定机制,实体归因验证可检测DG,现有框架未实施,为该领域研究提供新视角和方法。

Comments 24 pages, 7 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06748 2026-06-30 cs.CL cs.AI cs.LG 新提交 86%

Evidence Graph Consistency in Retrieval-Augmented Generation: A Model-Dependent Analysis of Hallucination Detection

检索增强生成中的证据图一致性:基于模型的幻觉检测分析

Jianru Shen

机构 * University of Montana(蒙大拿大学)

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 提出证据图一致性(EGC)框架,通过构建局部证据图并计算五种结构一致性指标检测幻觉,发现不同模型族间一致性特征方向相反,表明嵌入图一致性不能作为模型无关的检测信号。

Comments Accepted at the International Conference on Advanced Machine Learning and Data Science; to appear in the IEEE Xplore proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19493 2026-06-09 cs.CL cs.AI 86%

Development and Evaluation of a Retrieval-Augmented Generation Tool for Creating SAPPhIRE Models of Artificial Systems

SAPPhIRE人工系统模型创建工具的开发与评估

Anubhab Majumder, Kausik Bhattacharya, Amaresh Chakrabarti

机构 * Department of Design and Manufacturing, Indian Institute of Science(设计与制造系,印度科学研究院)

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文提出一种基于检索增强生成的工具,用于创建SAPPhIRE因果模型的人工系统模型,通过评估工具在事实准确性和可靠性方面的表现,提升系统设计类比支持能力。

Comments This paper has been accepted for presentation at the 10th International Conference on Research Into Design, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00957 2026-05-05 cs.IR cs.AI 86%

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation

我不知道--迈向具有确定性意识的适当信任

Daan Di Scala, Maaike de Boer, Pınar Yolum

机构 * TNO Netherlands Organisation for Applied Scientific Research, Department Data Science(荷兰应用科学研究院,数据科学部门) Utrecht University, Department of Information and Computing Sciences(乌得勒支大学,信息与计算科学系)

专题命中 RAG评测 :retrieval augmented generation(title,abstract);RAG(abstract,abstract_cn);分类 cs.IR、cs.AI

AI总结 本文提出CERTA系统,通过结合问题、上下文和答案的相关性来反映不确定性,以建立适当的信任。研究创建了Certainty Benchmark,并通过实验验证了CERTA在减少过度同意和提供谨慎行为方面的有效性。

Comments To be published in VALE 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24218 2026-03-26 cs.IR cs.AI 86%

Who Benefits from RAG? The Role of Exposure, Utility and Attribution Bias

谁从RAG中受益?曝光、效用和归因偏差的作用

Mahdi Dehghan, Graham McDonald

机构 * University of Glasgow, Glasgow, UK(格拉斯哥大学,格拉斯哥,英国)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.IR、cs.AI

AI总结 本文研究RAG中查询组公平性的影响因素,发现RAG系统在不同组查询的平均准确率和改进上存在偏差,揭示了曝光、效用和归因对公平性的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21553 2026-02-26 cs.IR cs.AI cs.LG 86%

Revisiting RAG Retrievers: An Information Theoretic Benchmark

重新审视RAG检索器:一个信息论基准

Wenqing Zheng, Dmitri Kalaev, Noah Fatsi, Daniel Barcklow, Owen Reinert, Igor Melnyk, Senthil Kumar, C. Bayan Bruss

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.IR、cs.AI

AI总结 本文提出MIGRASCOPE基准,通过信息论方法分析RAG检索器的性能、冗余和协同效应,揭示最佳检索器组合策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05863 2025-12-08 cs.CL cs.AI 86%

Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework

优化医疗问答系统:基于RAG框架的微调与零样本大语言模型比较研究

Tasnimul Hassan, Md Faisal Karim, Haziq Jeelani, Elham Behnam, Robert Green, Fayeq Jeelani Syed

机构 * Department of Electrical Engineering Computer Science University of Toledo Toledo, USA Institute of Mathematical Sciences Claremont Graduate University Claremont, USA Department of Bioengineering University of Toledo Toledo, USA Department of Computer Science Bowling Green State University Bowling Green, USA

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);knowledge retrieval(abstract);分类 cs.CL、cs.AI

AI总结 本文通过RAG框架结合微调与零样本大语言模型,提升医疗问答系统的准确性与可靠性,实验证明检索增强显著提高回答质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14412 2025-08-13 cs.IR cs.AI cs.LG 86%

RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition

Tim Cofala, Oleh Astappiev, William Xion, Hailay Teklehaymanot

机构 * L3S Research Center(L3S研究所以) Inria Paris-Rocquencourt(Inria巴黎-罗克琴克特研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒姆研究实验室)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.IR、cs.AI

Comments 4 pages, 6 figures. Report for SIGIR 2025 LiveRAG Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01954 2025-06-03 cs.CL cs.AI cs.LG 86%

DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation

Jennifer Chen, Aidar Myrzakhan, Yaxin Luo, Hassaan Muhammad Khan, Sondos Mahmoud Bsharat, Zhiqiang Shen

机构 * VILA Lab(VILA实验室) Mohamed bin Zayed University of AI(莫扎伊德·本·泽德人工智能大学) McGill University(麦吉尔大学) National University of Science and Technology(国家科学与技术大学)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);knowledge retrieval(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 Main. Code is available at https://github.com/VILA-Lab/DRAG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17137 2025-04-25 cs.CL cs.AI 86%

MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation

Chanhee Park, Hyeonseok Moon, Chanjun Park, Heuiseok Lim

机构 * Korea University, Republic of Korea(韩国大学)

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);retriever(abstract);分类 cs.CL、cs.AI

Comments Accepted to NAACL2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11005 2025-01-17 cs.CL cs.AI 86%

RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems

Robert Friel, Masha Belyi, Atindriyo Sanyal

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);retriever(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10571 2024-12-24 cs.CL cs.IR 86%

Evidence Contextualization and Counterfactual Attribution for Conversational QA over Heterogeneous Data with RAG Systems

Rishiraj Saha Roy, Joel Schlotthauer, Chris Hinze, Andreas Foltyn, Luzian Hahn, Fabian Kuech

专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);retriever(abstract);分类 cs.IR、cs.CL

Comments Accepted at WSDM 2025, 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06575 2024-06-12 cs.CL cs.AI 86%

Ask-EDA: A Design Assistant Empowered by LLM, Hybrid RAG and Abbreviation De-hallucination

Luyao Shi, Michael Kazda, Bradley Sears, Nick Shropshire, Ruchir Puri

专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);hybrid retrieval(abstract);分类 cs.CL、cs.AI

Comments Accepted paper at The First IEEE International Workshop on LLM-Aided Design, 2024 (LAD 24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11267 2026-07-14 cs.IR cs.AI cs.CL 新提交 86%

Enhancing LLMs through human feedback: a journey towards self-improvement

通过人类反馈增强语言模型:自我提升之旅

Tatiana Pelc, Gila Kamhi, Asaf Avrahamy, Adi Fledel-Alon

机构 * Intel Corporation(英特尔公司)

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 研究在信息检索系统中,通过整合辅助反馈RAG系统及人工参与,利用人类反馈优化主RAG系统性能,经多数据集测试验证方法有效,强调了其变革潜力,为自适应信息检索技术研究树立了先例。

Comments AIC 2025: The 10th International Workshop on Artificial Intelligence and Cognition (held as part of ECAI 2025). October 25-26, 2025. Bologna, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24200 2026-06-24 cs.CL cs.AI cs.IR 新提交 86%

MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval

MMed-Bench-IR:多语言医学信息检索的异构基准

Junhyeok Lee, Han Jang, Hyeonjin Goh, Kyu Sung Choi

机构 * Seoul National University(首尔国立大学) Seoul National University College of Medicine(首尔国立大学医学院) Seoul National University Hospital(首尔国立大学医院)

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 针对临床RAG的多语言检索需求,提出MMed-Bench-IR基准,涵盖跨语言对齐、概念区分和证据检索三种异构任务,评估6种语言下10个系统,发现生物医学编码器在英语与日语间性能差距巨大。

Comments Under review. 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26184 2026-04-30 cs.IR cs.AI cs.CL 86%

Auto-ARGUE: LLM-Based Report Generation Evaluation

Auto-ARGUE:基于大语言模型的报告生成评估

William Walden, Marc Mason, Orion Weller, Laura Dietz, John Conroy, Neil Molino, Hannah Recknor, Bryan Li, Gabrielle Kaili-May Liu, Yu Hou, Dawn Lawrie, James Mayfield, Eugene Yang

机构 * Johns Hopkins University(约翰霍普金斯大学) University of New Hampshire(新罕布什尔大学) IDA Center for Computing Services(IDA计算服务中心) Google(谷歌) Yale University(耶鲁大学) University of Maryland(马里兰大学)

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 本文提出Auto-ARGUE,一种基于大语言模型的报告生成评估工具,通过TREC 2024 NeuCLIR和RAG任务验证,展示了与人类判断的良好相关性,并发布ARGUE-Viz可视化工具。

Comments SIGIR 2026: Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13141 2026-06-12 cs.AI 新提交 85%

Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

重新思考长视频中的RAG:检索什么以及如何使用?

Yuho Lee, Jisu Shin, Nicole Hee-Yeon Kim, Jihwan Bang, Juntae Lee, Kyuwoong Hwang, Fatih Porikli, Hwanjun Song

机构 * Department of Computer Science, Cranberry-Lemon University(蔓越莓柠檬大学计算机科学系)

专题命中 RAG评测 :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 针对视频检索增强生成中检索粒度单一和基准测试缺陷,提出V-RAGBench基准和CARVE方法,通过分块自适应重排序实现多配置交错证据,显著提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07612 2026-03-10 cs.CL 85%

KohakuRAG: A simple RAG framework with hierarchical document indexing

KohakuRAG:一种具有层次文档索引的简单RAG框架

Shih-Ying Yeh, Yueh-Feng Ku, Ko-Wei Huang, Buu-Khang Tu

机构 * National Tsing Hua University Comfy Org Research Kohaku-Lab

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);dense retrieval(abstract);分类 cs.CL

AI总结 KohakuRAG通过层次文档索引和集合推理技术,提升了检索增强生成系统在高精度引用任务中的性能,取得挑战赛第一名。

Comments 38pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19927 2026-01-29 cs.CL 85%

Attribution Techniques for Mitigating Hallucinated Information in RAG Systems: A Survey

缓解RAG系统中幻觉信息的归因技术:综述

Yuqing Zhao, Ziyao Liu, Yongsen Zheng, Kwok-Yan Lam

机构 * Nanyang Technological University(南洋理工大学)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.CL

AI总结 本文综述了RAG系统中缓解幻觉的归因技术,通过分类幻觉类型、统一流程和比较优劣,为实际应用提供指导。

Journal ref The 8th International Conference on Artifcial Intelligence in Information and Communication (ICAIIC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02932 2024-10-07 cs.AI 85%

Intrinsic Evaluation of RAG Systems for Deep-Logic Questions

Junyi Hu, You Zhou, Jie Wang

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04369 2024-06-10 cs.SE cs.AI 85%

RAG Does Not Work for Enterprises

Tilmann Bruckhaus

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);knowledge retrieval(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20735 2026-02-25 cs.IR cs.AI cs.CL 85%

RMIT-ADM+S at the MMU-RAG NeurIPS 2025 Competition

RMIT-ADM+S在MMU-RAG NeurIPS 2025比赛中的表现

Kun Ran, Marwah Alaofi, Danula Hettiachchi, Chenglong Ma, Khoi Nguyen Dinh Anh, Khoi Vo Nguyen, Sachin Pathiyan Cherumanal, Lida Rashidi, Falk Scholer, Damiano Spina, Shuoqi Sun, Oleg Zendel

机构 * RMIT University(皇家墨尔本理工大学)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 RMIT-ADM+S系统通过Routing-to-RAG架构在NeurIPS 2025 MMU-RAG比赛中获胜,实现了高效检索增强生成和复杂科研任务处理。

Comments MMU-RAG NeurIPS 2025 winning system

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14400 2026-07-17 cs.CL cs.IR 新提交 85%

DS@GT ARC at LongEval: Citation Integrity and Factual Grounding in Scientific QA

DS@GT ARC参加LongEval:科学问答中的引用完整性和事实基础

Brandon Michaels, Brendon Johnson

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

AI总结 研究在科学问答中传统评估指标与引用完整性的差异,通过Corrective RAG和CiteFix构建纠正管道,对比前沿模型,发现前沿模型答案生成不依赖文档上下文,而纠正管道提升了引用忠实度和答案基础,提出需奖励严格答案基础的评估指标。

Comments 12 pages, 4 figures. Accepted to the CLEF 2026 LongEval Lab Working Notes

详情

展开后加载摘要…

URL PDF HTML 收藏