arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 1199 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 1199 篇

2606.15646 2026-06-16 cs.AI 新提交 84%

NeuroSymbolic AI for Legal AI-TRISM: Trustworthy, Reliable, Interpretable, Safe Models

面向法律AI-TRISM的神经符号AI:可信、可靠、可解释、安全模型

Deepa Tilwani, Yash Saxena, Ankur Padia, Srinivasan Parthasarathy, Manas Gaur

机构 * Department of Computer Science, AI Institute, University of South Carolina(南卡罗来纳大学计算机科学系,人工智能研究所) Department of Computer Science and Electrical Engineering, University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校计算机科学与电气工程系) Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系)

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 针对法律领域LLM缺乏可解释推理和易产生幻觉的问题,提出TRISM框架,融合神经符号AI与LLM,通过结构化法律知识集成和RAG验证机制提升模型可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22827 2026-06-16 cs.SE cs.AI 版本更新 84%

FasterPy: An LLM-based Code Execution Efficiency Optimization Framework

FasterPy:基于大语言模型的代码执行效率优化框架

Yue Wu, Minghao Han, Ruiyin Li, Peng Liang, Amjed Tahir, Zengyang Li, Qiong Feng, Mojtaba Shahin

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机学院) School of Mathematical and Computational Sciences, Massey University(梅西大学数学与计算科学学院) School of Computer Science, Central China Normal University(中央中国师范大学计算机学院) School of Computer Science, Nanjing University of Science and Technology(南京理工大学计算机学院) School of Computing Technologies, RMIT University(皇家墨尔本理工大学计算技术学院)

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出FasterPy框架,结合检索增强生成(RAG)和低秩适应(LoRA)技术,利用大语言模型自动优化Python代码执行效率,在PIE基准上超越现有方法。

Comments 38 pages, 5 images, 14 tables, Manuscript revision submitted to a Journal (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14119 2026-06-15 cs.AI 新提交 84%

FactoryLLM: A Safe and Open-Source AI Playground for Evaluating LLMs in Smart Factories

FactoryLLM:用于评估智能工厂中大语言模型的安全开源AI实验场

Yash Pulse, Yong-Bin Kang, Abhik Banerjee, Abdur Forkan, Prem Prakash Jayaraman

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出FactoryLLM,一个安全开源的AI实验场,通过多机器文档分析评估基于RAG的大语言模型,采用RAGAS和NVIDIA LLM-as-a-Judge双评估机制,案例验证了跨机器文档推理的有效性。

Comments 6 pages, 3 figures, IEEE INDIN 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12897 2026-06-12 cs.CL 新提交 84%

SafeLLM: Extraction as a Hallucination-Resistant Alternative to Rewriting in Safety-Critical Settings

SafeLLM: 在安全关键场景中,提取作为重写的抗幻觉替代方案

Julia Ive, Felix Jozsa, Evridiki Georgaki, Nabeel Sheikh, Emma Cattell, Nick Jackson, Paulina Bondaronek, Ciaran Scott Hill, Richard Dobson

机构 * Institute of Health Informatics, University College London(伦敦大学学院健康信息学研究所) National Hospital for Neurology and Neurosurgery(国家神经内科与神经外科医院) Somerset NHS Foundation Trust(萨默塞特NHS基金会信托) King's College Hospital(国王学院医院) King's College London(伦敦国王学院)

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.CL

AI总结 提出将提取作为重写型RAG的抗幻觉替代方案,通过行号选择策略在安全关键文档中实现高召回(95%)和低幻觉,优于直接复制和安全导向方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19743 2026-05-28 cs.AI cs.LG cs.MA 84%

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

EngiAI: 面向LLM驱动工程设计的智能体框架与基准测试套件

Gioele Molinari, Florian Felten, Soheyl Massoudi, Mark Fuge

机构 * IDEAL Chair of Artificial Intelligence in Engineering Design(人工智能与工程设计理想 chair) ETH Zurich(苏黎世联邦理工学院) Autom8.build

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出EngiAI多智能体系统框架和包含工作流、RAG、HPC三维度的基准套件,通过监督架构协调七个专业智能体,验证了LLM在工程设计中的能力与局限。

Comments 26 pages, 10 figures, to be published at IDETC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10177 2026-04-24 cs.AI 84%

KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for Machine Learning Problems

KompeteAI:面向机器学习问题端到端流水线生成的加速自主多智能体系统

Stepan Kulibaba, Artem Dzhalilov, Roman Pakhomov, Oleg Svidchenko, Alexander Gasnikov, Aleksei Shpilman

机构 * Research Center of the Artificial Intelligence Institute(人工智能研究所研究中心) Innopolis University(因诺波利斯大学) Sberbank of Russia(俄罗斯Sberbank) AI4S Center(AI4S中心) MIPT(莫斯科国立信息安全大学) Steklov Institute(斯捷克洛夫研究所)

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 KompeteAI通过动态探索解决方案空间和整合RAG技术,提升了AutoML的探索效率与执行速度,实现6.9倍的流水线评估加速,并在MLE-Bench基准上超越主流方法3%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03739 2026-07-07 cs.CR cs.AI cs.CL 新提交 84%

A Failure-Mode Benchmark for Polymorphic Sybil Poisoning in RAG

RAG中多态性女巫中毒的故障模式基准测试

Donghyun Lee, Juntae Kim

机构 * Department of Computer Engineering(计算机工程系) Dongguk University(东国大学)

专题命中 RAG评测 :RAG(title,title_cn);分类 cs.CL、cs.AI

AI总结 该研究发布了用于问答的基准测试和评估框架,划分读者输出类别,介绍多态性女巫中毒攻击,通过对比揭示其危害,还发布了冻结基准、评估工具等。

Comments 17 pages, 2 figures, 11 tables. Corresponding author: Juntae Kim. Dataset and code to be released upon publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10874 2026-04-14 cs.CL cs.AI 84%

AOP-Smart: A RAG-Enhanced Large Language Model Framework for Adverse Outcome Pathway Analysis

AOP-Smart:一种增强检索的大型语言模型框架用于不良后果路径分析

Qinjiang Niu, Lu Yan

机构 * Nanyang Normal University(南阳师范学院)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出AOP-Smart框架,通过检索增强生成技术提升大型语言模型在不良后果路径分析中的可靠性与准确性,实验表明其显著缓解了模型幻觉问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21359 2026-03-24 cs.CL cs.AI cs.CY 84%

Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF

印地语方言偏见评估:整合基于RAG的翻译和人类增强的RLAIF的多阶段框架

K. M. Jubair Sami, Dipto Sumit, Ariyan Hossain, Farig Sadeque

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一个多阶段框架,通过基于RAG的翻译和人类增强的RLAIF评估九种印地语方言的偏见。研究发现方言差异导致性能下降,模型规模增加不总是缓解偏见。

Comments 12 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01183 2026-03-23 cs.CL cs.AI 84%

TempPerturb-Eval: On the Joint Effects of Internal Temperature and External Perturbations in RAG Robustness

TempPerturb-Eval: 关于RAG鲁棒性中内部温度与外部扰动的联合影响

Yongxin Zhou, Philippe Mulhem, Didier Schwab

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了RAG系统中内部温度与外部扰动的交互影响,提出分析框架并展示实验结果,揭示高温度设置加剧扰动敏感性,提出评估基准和调参指南。

Comments LREC 2026, Palma, Mallorca (Spain), 11-16 May 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13538 2026-03-19 cs.IR cs.AI 84%

RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines

RAGXplain: 从可解释评估到RAG流水线的可操作指导

Dvir Cohen, Tamir Houri, Lin Burg, Gilad Barkan

机构 * Wix.com AI Research(Wix.com人工智能研究)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 RAGXplain通过'指标钻石'框架将性能指标转化为可操作指导,利用LLM生成自然语言失败模式解释,提升RAG流水线性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14173 2026-03-17 cs.LG cs.AI cs.IR 84%

Hybrid Intent-Aware Personalization with Machine Learning and RAG-Enabled Large Language Models for Financial Services Marketing

混合意图感知个性化:结合机器学习和RAG增强的大语言模型用于金融服务营销

Akhil Chandra Shanivendra

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文提出一种混合架构,结合传统机器学习与RAG增强的大语言模型,用于金融服务营销中的个性化营销,通过时间编码器、潜在表示和多任务分类提高个性化准确性。

Comments 18 pages, 5 figures, 3 tables. Applied ML systems paper. The contribution is architectural rather than algorithmic

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04406 2026-03-06 cs.CL cs.AI 84%

CTRL-RAG: Contrastive Likelihood Reward Based Reinforcement Learning for Context-Faithful RAG Models

CTRL-RAG:基于对比似然奖励的强化学习用于上下文忠实的RAG模型

Zhehao Tan, Yihan Jiao, Dan Yang, Junjie Wang, Duolin Sun, Jie Feng, Xidong Wang, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 CTRL-RAG通过对比似然奖励机制提升RAG模型的上下文忠实性,结合内部与外部奖励框架,在多个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09886 2026-02-26 cs.CL cs.AI 84%

Probabilistic distances-based hallucination detection in LLMs with RAG

基于概率距离的LLM中幻觉检测方法(RAG)

Rodion Oblovatny, Alexandra Kuleshova, Konstantin Polev, Alexey Zaytsev

机构 * Markov Lab, Department of Mathematics(马尔可夫实验室,数学系) Computer Science, Saint-Petersburg University(计算机科学,圣彼得堡大学) AI Center, Skoltech(人工智能中心,斯克里普丘克技术学院) SB AI Lab(SB人工智能实验室) AI Center, Skoltech, Risk department, Sber(人工智能中心,斯克里普丘克技术学院,风险部门)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种基于概率距离的LLM幻觉检测方法,利用提示标记与响应标记嵌入分布的距离来检测幻觉,具有高效性和可迁移性。

Comments Updated approach to constructing a hallucination detection score. Added results from experiments with the NLI task. The approach with trainable deep kernels has been removed, with a focus on the unsupervised approach

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20379 2026-02-25 cs.CL cs.AI 84%

Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems

面向企业级RAG系统的案例感知LLM-as-a-Judge评估

Mukul Chhabra, Luigi Medrano, Arush Verma

机构 * Dell Technologies(戴尔技术)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出面向企业级RAG系统的案例感知LLM-as-a-Judge评估框架,通过八个操作指标评估多轮工作流,揭示企业关键权衡,提升系统诊断清晰度和可操作性。

Comments 12 pages including appendix, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17856 2026-02-23 cs.IR cs.AI 84%

Enhancing Scientific Literature Chatbots with Retrieval-Augmented Generation: A Performance Evaluation of Vector and Graph-Based Systems

通过检索增强生成增强科学文献聊天机器人:向量和图基系统性能评估

Hamideh Ghanadian, Amin Kamali, Mohammad Hossein Tekieh

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.IR、cs.AI

AI总结 本文通过对比向量和图基检索系统,评估了检索增强生成在提升科学文献聊天机器人性能中的效果,展示了混合系统在提高科学知识可及性方面的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16415 2026-02-12 cs.CL cs.AI cs.LG 84%

Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation

基于Jensen-Shannon散度的响应归因研究:一种驱动上下文归因的机理研究

Ruizhe Li, Chen Chen, Yuchen Hu, Yanjun Gao, Xi Wang, Emine Yilmaz

机构 * University of Aberdeen(阿伯丁大学) Nanyang Technological University(南洋理工大学) University College London(伦敦大学学院) University of Colorado Anschutz Medical Campus(科罗拉多大学安舒兹医学校区) University of Sheffield(谢菲尔德大学)

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

AI总结 本文提出基于Jensen-Shannon散度的ARC-JSD方法,实现高效准确的上下文归因,提升RAG模型的性能和效率。

Comments Accepted at ICLR 2026; Best Paper Award at COLM 2025 XLLM-Reason-Plan Workshop; Accepted at NeurIPS 2025 Mechanistic Interpretability Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11847 2026-02-12 cs.IR cs.AI cs.CY 84%

A Multimodal Manufacturing Safety Chatbot: Knowledge Base Design, Benchmark Development, and Evaluation of Multiple RAG Approaches

多模态制造安全聊天机器人:知识库设计、基准开发及多种RAG方法的评估

Ryan Singh, Austin Hamilton, Amanda White, Michael Wise, Ibrahim Yousif, Arthur Carvalho, Zhe Shan, Reza Abrisham Baf, Mohammad Mayyas, Lora A. Cavuoto, Fadel M. Megahed

机构 * Farmer School of Business, Miami University(Miami大学农业商学院) Department of Computer Science and Software Engineering, Miami University(Miami大学计算机科学与软件工程系) Department of Engineering Technology, Miami University(Miami大学工程技术系) Department of Mechanical and Manufacturing Engineering, Miami University(Miami大学机械与制造工程系) Department of Industrial and Systems Engineering, University at Buffalo(University at Buffalo工业与系统工程系)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文提出了一种多模态制造安全聊天机器人,通过RAG方法实现高准确率、低延迟和低成本的安全培训,展示了其在工业5.0环境中的应用价值。

Comments 25 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09442 2026-02-11 cs.CL cs.AI 84%

Evaluating Social Bias in RAG Systems: When External Context Helps and Reasoning Hurts

评估RAG系统中的社会偏见:当外部上下文有助于而推理有害

Shweta Parihar, Lu Cheng

机构 * University of Illinois at Chicago(伊利诺伊大学芝加哥分校)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本研究评估RAG系统中的社会偏见,发现外部上下文可减少偏见,但链式思维提示反而增加偏见,凸显了需要偏见意识推理框架的重要性。

Comments Accepted as a full paper with an oral presentation at the 30th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15487 2026-01-23 cs.AI cs.CL cs.MA 84%

MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation

MiRAGE:一种多智能体框架,用于生成多模态多跳问题-答案数据集以评估RAG系统

Chandan Kumar Sahu, Premith Kumar Chilukuri, Matthew Hetrich

机构 * ABB Inc(ABB公司)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 MiRAGE通过多智能体框架生成多模态多跳问题-答案数据集,提升RAG系统评估的准确性和复杂性。

Comments 12 pages, 2 figures, Submitted to ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16101 2026-01-21 cs.AI cs.IR 84%

Worse than Zero-shot? A Fact-Checking Dataset for Evaluating the Robustness of RAG Against Misleading Retrievals

比零样本更差?一个评估RAG在误导检索中鲁棒性的事实核查数据集

Linda Zeng, Rithwik Gupta, Divij Motwani, Yi Zhang, Diji Yang

机构 * The Harker School(哈克尔学校) Irvington High School(艾尔文顿高中) Palo Alto High School(帕洛阿尔托高中) University of California Santa Cruz(加州大学圣克鲁兹分校)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 RAGuard是首个评估RAG系统在误导检索中鲁棒性的基准,揭示LLM在面对误导信息时的表现劣于零样本基线。

Comments Advances in Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04196 2026-01-09 cs.CL cs.IR 84%

RAGVUE: A Diagnostic View for Explainable and Automated Evaluation of Retrieval-Augmented Generation

RAGVUE:一种用于可解释和自动评估检索增强生成的诊断视图

Keerthana Murugaraj, Salima Lamsiyah, Martin Theobald

机构 * University of Luxembourg(卢森堡大学)

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.IR、cs.CL

AI总结 RAGVUE通过分解RAG行为并提供结构化解释,实现无参考的自动评估,揭示RAG流程中的细粒度失败。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15068 2025-12-22 cs.LG cs.AI cs.CL 84%

The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems

语义幻觉:基于嵌入的幻觉检测在RAG系统中的认证极限

Debu Sinha

机构 * Independent Researcher(独立研究者)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 研究揭示了基于嵌入的幻觉检测在RAG系统中存在语义幻觉问题,通过符合预测方法发现真实幻觉检测的挑战,证明需通过推理而非表面语义解决。

Comments 12 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12524 2025-12-19 cs.CL cs.AI 84%

Enhancing Long-term RAG Chatbots with Psychological Models of Memory Importance and Forgetting

通过记忆重要性与遗忘的心理模型增强长期RAG聊天机器人

Ryuichi Sumida, Koji Inoue, Tatsuya Kawahara

机构 * Graduate School of Informatics Kyoto University(京都大学信息学研究科)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 LUFY通过优先保留情绪唤醒记忆并遗忘大部分对话内容,提升长期对话体验和聊天机器人性能。

Comments 37 pages, accepted and published in Dialogue & Discourse 16(2) (2025)

Journal ref Dialogue & Discourse 16(2) (2025) 74--110

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01335 2025-12-02 cs.CR cs.AI cs.CL 84%

EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations

EmoRAG:评估RAG对符号扰动的鲁棒性

Xinyun Zhou, Xinfeng Li, Yinan Peng, Ming Xu, Xuanwang Zhang, Miao Yu, Yidong Wang, Xiaojun Jia, Kun Wang, Qingsong Wen, XiaoFeng Wang, Wei Dong

机构 * ZJU Hangzhou China(浙江大学杭州校区) NTU Singapore(南洋理工大学) Hengxin Tech. Singapore(新加坡恒心科技) NUS Singapore(国立新加坡大学) NJU Nanjing China(南京大学) PKU Beijing China(北京大学)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 EmoRAG研究揭示RAG系统对细微表情符号扰动的鲁棒性问题,发现单个表情符号可导致检索严重误导,并提出针对性防御措施。

Comments Accepted to ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14531 2025-11-19 cs.CL cs.IR 84%

LiveRAG: A diverse Q&A dataset with varying difficulty level for RAG evaluation

David Carmel, Simone Filice, Guy Horowitz, Yoelle Maarek, Alex Shtoff, Oren Somekh, Ran Tavory

机构 * Technology Innovation Institute (TII)(技术创新研究所)

专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.IR、cs.CL

Comments 14 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14066 2025-11-19 cs.IR cs.AI 84%

Retrieval-Augmented Generation in Industry: An Interview Study on Use Cases, Requirements, Challenges, and Evaluation

Lorenz Brehme, Benedikt Dornauer, Thomas Ströhle, Maximilian Ehrhart, Ruth Breu

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.IR、cs.AI

Comments This preprint was accepted for presentation at the 17th International Conference on Knowledge Discovery and Information Retrieval (KDIR25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04502 2025-11-07 cs.CL cs.AI 84%

RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG

Joshua Gao, Quoc Huy Pham, Subin Varghese, Silwal Saurav, Vedhus Hoskere

机构 * University of Houston(德克萨斯大学休斯敦分校)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04847 2025-11-07 cs.CL cs.AI 84%

Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards

Manveer Singh Tamber, Forrest Sheng Bao, Chenyu Xu, Ge Luo, Suleman Kazi, Minseok Bae, Miaoran Li, Ofer Mendelevitch, Renyi Qu, Jimmy Lin

机构 * University of Waterloo(滑铁卢大学) Vectara(Vectara公司) Iowa State University(爱荷华州立大学) Stanford University(斯坦福大学)

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

Comments EMNLP Industry Track 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03261 2025-11-06 cs.CL cs.AI 84%

Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature

Ranul Dayarathne, Uvini Ranaweera, Upeksha Ganegoda

机构 * University of Moratuwa(穆塔瓦大学)

专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.CL、cs.AI

Comments 18 pages, 4 figures, 5 tables, presented at the 5th International Conference on Artificial Intelligence in Education Technology

Journal ref Lecture Notes on Data Engineering and Communications Technologies, vol. 228, Springer, 2025, pp. 387--403

详情

展开后加载摘要…

URL PDF HTML 收藏