arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 1199 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 1199 篇

2502.15854 2025-02-25 cs.LG cs.AI cs.CL 84%

Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models

Aryan Jadon, Avinash Patil, Shashank Kumar

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

Comments 8 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02850 2025-02-19 cs.CY cs.AI cs.HC cs.IR 84%

WASHtsApp -- A RAG-powered WhatsApp Chatbot for supporting rural African clean water access, sanitation and hygiene

Simon Kloker, Alex Cedric Luyima, Matthew Bazanya

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

Comments Working Paper. Accepted at IST-Africa Conference 2025, Nairobi

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07437 2025-02-17 cs.CL cs.AI 84%

Evaluation of Retrieval-Augmented Generation: A Survey

Hao Yu, Aoran Gan, Kai Zhang, Shiwei Tong, Qi Liu, Zhaofeng Liu

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12101 2025-02-14 cs.CL cs.AI 84%

Better RAG using Relevant Information Gain

Marc Pickett, Jeremy Hartman, Ayan Kumar Bhowmick, Raquib-ul Alam, Aditya Vempaty

专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.CL、cs.AI

Comments 4 page paper submitted to EMNLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12789 2025-01-23 cs.CL cs.IR 84%

Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana

Simone Filice, Guy Horowitz, David Carmel, Zohar Karnin, Liane Lewin-Eytan, Yoelle Maarek

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13720 2025-01-09 cs.CL cs.AI 84%

Federated Learning and RAG Integration: A Scalable Approach for Medical Large Language Models

Jincheol Jung, Hongju Jeong, Eui-Nam Huh

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03468 2025-01-08 cs.CL cs.AI 84%

MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems

Yannis Katsis, Sara Rosenthal, Kshitij Fadnis, Chulaka Gunasekara, Young-Suk Lee, Lucian Popa, Vraj Shah, Huaiyu Zhu, Danish Contractor, Marina Danilevsky

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14510 2024-12-20 cs.CL cs.AI 84%

PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization

Jiayi Wu, Hengyi Cai, Lingyong Yan, Hao Sun, Xiang Li, Shuaiqiang Wang, Dawei Yin, Ming Gao

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10151 2024-12-16 cs.CV cs.AI cs.CL 84%

VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation

Hyeonseok Lim, Dongjae Shin, Seohyun Song, Inho Won, Minjun Kim, Junghun Yuk, Haneol Jang, KyungTae Lim

专题命中 RAG评测 :retrieval augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

Comments The 31st International Conference on Computational Linguistics (COLING 2025), 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13663 2024-12-02 cs.CL cs.AI cs.LG 84%

Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation

Jirui Qi, Gabriele Sarti, Raquel Fernández, Arianna Bisazza

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2024 Main Conference. Code and data released at https://github.com/Betswish/MIRAGE

Journal ref Proceedings of EMNLP (2024) 6037-6053

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09607 2024-11-15 cs.IR cs.CL 84%

Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework

Ronak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos, Nick Craswell, Jimmy Lin

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15187 2024-11-01 cs.AI cs.IR 84%

UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis

Yulong Hui, Yao Lu, Huanchen Zhang

专题命中 RAG评测 :retrieval augmented generation(title);retrieval-augmented generation(abstract);RAG(abstract);分类 cs.IR、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15763 2024-09-27 cs.IR cs.AI 84%

IRSC: A Zero-shot Evaluation Benchmark for Information Retrieval through Semantic Comprehension in Retrieval-Augmented Generation Scenarios

Hai Lin, Shaoxiong Zhan, Junyou Su, Haitao Zheng, Hui Wang

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.IR、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09916 2024-09-17 cs.CL cs.AI 84%

SFR-RAG: Towards Contextually Faithful LLMs

Xuan-Phi Nguyen, Shrey Pandit, Senthil Purushwalkam, Austin Xu, Hailin Chen, Yifei Ming, Zixuan Ke, Silvio Savarese, Caiming Xong, Shafiq Joty

专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.CL、cs.AI

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17587 2024-08-19 cs.IR cs.CL cs.LG 84%

RAGSys: Item-Cold-Start Recommender as RAG System

Emile Contal, Garrin McGoldrick

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12057 2024-07-18 cs.CL cs.AI 84%

NinjaLLM: Fast, Scalable and Cost-effective RAG using Amazon SageMaker and AWS Trainium and Inferentia2

Tengfei Xue, Xuefeng Li, Roman Smirnov, Tahir Azim, Arash Sadrieh, Babak Pahlavan

专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01080 2024-07-10 cs.CL cs.AI 84%

Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese

Yunqi Xu, Tianchi Cai, Jiyan Jiang, Xierui Song

专题命中 RAG评测 :retrieval augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

Journal ref KDD 2024 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05654 2024-06-18 cs.CL cs.IR 84%

DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation

Shuting Wang, Jiongnan Liu, Shiren Song, Jiehan Cheng, Yuqi Fu, Peidong Guo, Kun Fang, Yutao Zhu, Zhicheng Dou

专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.IR、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05444 2024-05-10 cs.CL cs.AI 84%

Evaluating Students' Open-ended Written Responses with LLMs: Using the RAG Framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large

Jussi S. Jauhiainen, Agustín Garagorry Guerra

专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.CL、cs.AI

Comments 18 pages, 6 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00820 2024-03-05 cs.IR cs.CL 84%

Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup

Tristan Kenneweg, Philip Kenneweg, Barbara Hammer

专题命中 RAG评测 :retrieval augmented generation(title,abstract);RAG(abstract);分类 cs.IR、cs.CL

Comments Was handed in to IJCNN prior to preprint publication here. Was neither accepted nor rejected at date of publication here

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00824 2026-08-04 cs.SE 新提交 83%

Structure-Aware Semantic Chunking with Title-Chain Prefixes: A 1600-Query Evaluation and the Measurement Trap in Text-Transform Ablations

结合标题链前缀的结构感知语义分块:1600次查询评估及文本变换消融实验中的测量陷阱

Yang Yang

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 本文提出仅在分块侧的三阶段语义分块流水线,经1600次查询评估提升RAG的MRR@5,还发现分块研究中存在的测量陷阱并推荐检索时每个候选前缀评估的方向。

Comments 8 pages, 1 table. Replication package: DOI 10.5281/zenodo.21744653

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19047 2026-07-01 cs.CL cs.AI cs.IR 版本更新 83%

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

RARE:面向高相似性语料的冗余感知检索评估框架

Hanjun Cho, Jay-Yoon Lee

机构 * Allganize Seoul National University(首尔国立大学)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 RARE通过分解文档为原子事实和增强LLM生成数据,构建更真实的基准,揭示当前评估体系无法捕捉的鲁棒性差距。

Comments Accepted to ACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25361 2026-06-25 cs.CL cs.AI cs.IR 新提交 83%

Memory Makes the Difference: Evaluating How Different Memory Roles Shape Conversational Agents

记忆决定差异:评估不同记忆角色如何塑造对话代理

Yuxin Wang, Paul Thomas, Zhiwei Yu, Yuan Gao, Saeed Hassanpour, Soroush Vosoughi, Robert Sim, Nick Craswell

机构 * Dartmouth College(达特茅斯学院) Microsoft(微软公司)

专题命中 RAG评测 :RAG(summary_cn,abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 本研究提出对话记忆细粒度分类法,通过用户中心评估框架,揭示不同类型记忆(如澄清记忆、无关记忆)对RAG对话系统响应准确性、个性化及主题相关性的差异化影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20359 2026-06-19 cs.LG 新提交 83%

Train, Retrieve, or Both? A Four-Arm Head-to-Head for Correct Statutory Citation on the Ontario Residential Tenancies Act

训练、检索,还是两者兼用?针对安大略省住宅租赁法的正确法定引用的四组头对头比较

Ali Asaria, Tony Salomone, Deep Gandhi

机构 * Transformer Lab

专题命中 RAG评测 :RAG(summary_cn,abstract);hybrid retrieval(abstract)

AI总结 研究自诉租户、房东和帮助台工作人员如何获得正确的法定引用,通过四组实验比较微调、检索及混合方法,发现SFT+RAG混合模型在精确匹配上得分最高且无幻觉引用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17179 2026-06-16 cs.MA 版本更新 83%

Ablation Study of a Fairness Auditing Agentic System for Bias Mitigation in Early-Onset Colorectal Cancer Detection

公平性审计代理系统在早发性结直肠癌检测中缓解偏见的消融研究

Amalia Ionescu, Jose Guadalupe Hernandez, Jui-Hsuan Chang, Emily F. Wong, Paul Wang, Jason H. Moore, Tiffani J. Bright

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 提出双代理架构(领域专家代理+公平性顾问代理),通过消融实验比较不同配置下大语言模型在公平性审计中的表现,发现带RAG的代理系统在识别差异方面语义相似度最高。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09845 2026-06-10 cs.HC cs.ET 新提交 83%

Tutor, Not Solver: Designing a Guardrailed AI Assistant for Learning in Higher Education: A Design Case of PeteChat

导师,而非解题者:设计高等教育中带护栏的AI学习助手——PeteChat的设计案例

Belle Li, Lily Tan, Wei Zakharov, Qiang Qiu, Colby Ben Acton

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 本文通过设计案例PeteChat,提出八项可迁移的评估感知AI导师设计原则,包括作业护栏、调试支架等,基于本地Llama-3模型和RAG技术,旨在平衡学习支持与学术诚信。

Comments Preprint. Includes supplementary appendices, interface figures, and baseline-analysis tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09552 2026-06-08 cs.IR cs.AI cs.CL 版本更新 83%

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval

MCERF:通过增强检索推进工程文档的多模态大语言模型评估

Kiarash Naghavi Khanghah, Hoang Anh Nguyen, Anna C. Doris, Amir Mohammad Vahedi, Daniele Grandi, Faez Ahmed, Hongyi Xu

机构 * School of Mechanical, Aerospace, and Manufacturing Engineering, University of Connecticut, Storrs, CT 06269(机械、航空航天与制造工程学院,康涅狄格大学,斯托尔斯,CT 06269) Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA(机械工程系,麻省理工学院,剑桥,MA 02139,美国)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval augmented generation(abstract);retriever(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 提出MCERF框架,结合多模态检索器ColPali与大语言模型推理,通过混合查找、视觉文本融合、高推理和自一致性决策等策略,在DesignQA基准上实现平均准确率相对提升41.1%,无需完整规则书摄入即可处理工程文档中的多模态问答。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05436 2026-06-05 cs.AI cs.CL cs.IR 83%

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison

十位头痛专家与人工智能在临床文献总结中的比较:一项关键评估与对比

Alejandro Lozano, Keiko Ihara, Ping-Hao Yang, Carrie E. Robertson, Jennifer Stern, Allan Purdy, Hsiangkuo Yuan, Pengfei Zhang, Yulia Orlova, Olga Fermo, Jennifer Hranilovich, Fred Cohen, Todd J. Schwedt, Jenelle A. Jindal, Serena Yeung-Levy, Chia-Chun Chiang

机构 * Stanford University Palo Alto CA USA(斯坦福大学) Department of Neurology Mayo Clinic Rochester MN USA(梅奥诊所神经科) Department of Neurology Dalhousie University Halifax Canada(达尔豪斯大学神经科) Jefferson Headache Center Department of Neurology Thomas Jefferson University PA USA(泰勒大学神经科) Beth Israel Deaconess Medical Center Boston MA USA(贝斯以色列医疗中心) Department of Neurology University of Florida Gainesville FL USA(佛罗里达大学神经科) University of Colorado School of Medicine Department of Pediatrics Division of Child Neurology Aurora CO USA(科罗拉多医学院儿科部儿童神经科) Department of Medicine Mount Sinai Hospital Icahn School of Medicine at Mount Sinai New York NY USA(西奈医院医学部) Department of Neurology Mayo Clinic Scottsdale AZ USA(梅奥诊所Scottsdale分部) Harvard Medical School Boston MA USA(哈佛医学院) Department of Neurology Mount Sinai Hospital Icahn School of Medicine at Mount Sinai New York NY USA(西奈医院神经科)

专题命中 RAG评测 :RAG(summary_cn,abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 本研究通过构建基于RAG的AI框架,比较了三种大语言模型与十位头痛专家在临床文献总结方面的表现,发现专家撰写的摘要更受青睐,但专家有时难以区分人类与AI生成的摘要。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11974 2026-05-13 cs.LG 83%

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization

迈向顺序公平性:通过双组优势优化缓解大语言模型的顺序敏感性

Xu Chu, Guanyu Wang, Zhijie Tan, Xinrong Chen, Ziyu Li, Tong Mo, Weiping Li

机构 * School of Software and Microelectronics, Peking University(软件与微电子学院,北京大学)

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 本文提出双组优势优化(DGAO),通过平衡组内准确性优势和组间稳定性优势,缓解LLM的顺序敏感性,提升模型在RAG、数学推理和分类任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12400 2026-04-16 cs.NI 83%

Agentic AI for 6G: A New Paradigm for Autonomous RAN Security Compliance

面向6G的代理AI:自主RAN安全合规的新范式

Sotiris Chatzimiltis, Mahdi Boloursaz Mashhadi, Mohammad Shojafar, Merouane Debbah, Rahim Tafazolli

专题命中 RAG评测 :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 本文提出利用基于LLM的AI代理与RAG流程框架,实现RAN安全合规的智能自主执行,通过案例研究展示其在O-RAN联盟和3GPP标准合规性评估中的应用,并探讨模型幻觉、供应商不一致等挑战及未来发展方向。

Comments Accepted at IEEE Communications Standards Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏