arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 1199 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 1199 篇

2511.01454 2025-11-04 cs.CL cs.DL 79%

"Don't Teach Minerva": Guiding LLMs Through Complex Syntax for Faithful Latin Translation with RAG

Sergio Torres Aguilar

专题命中 RAG评测 :RAG(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17514 2025-09-18 cs.AI 79%

TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment

Athanasios Davvetas, Xenia Ziouvelou, Ypatia Dami, Alexios Kaponis, Konstantina Giouvanopoulou, Michael Papademas

专题命中 RAG评测 :RAG(title,abstract);分类 cs.AI

Comments 9 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04026 2025-07-08 cs.CL 79%

Patient-Centered RAG for Oncology Visit Aid Following the Ottawa Decision Guide

Siyang Liu, Lawrence Chin-I An, Rada Mihalcea

机构 * The LIT Group, Department of Computer Science and Engineering, University of Michigan, Ann Arbor(密歇根大学计算机科学与工程系、LIT集团、安阿伯分校)

专题命中 RAG评测 :RAG(title);retrieval-augmented generation(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15911 2025-06-24 cs.CL 79%

From RAG to Agentic: Validating Islamic-Medicine Responses with LLM Agents

Mohammad Amaan Sayeed, Mohammed Talha Alam, Raza Imam, Shahab Saquib Sohail, Amir Hussain

机构 * Mohamed bin Zayed University of Artificial Intelligence, UAE(阿布扎克穆罕默德·本·扎耶德人工智能大学) Edinburgh Napier University, UK(爱丁堡纳皮尔大学) VIT Bhopal University, India(比哈尔大学)

专题命中 RAG评测 :RAG(title);retrieval-augmented generation(abstract);分类 cs.CL

Comments Published at the 4th Muslims in Machine Learning (MusIML) Workshop (ICML-25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01171 2025-06-24 cs.CL 79%

Multilingual Retrieval Augmented Generation for Culturally-Sensitive Tasks: A Benchmark for Cross-lingual Robustness

Bryan Li, Fiona Luo, Samar Haider, Adwait Agashe, Tammy Li, Runqi Liu, Muqing Miao, Shriya Ramakrishnan, Yuan Yuan, Chris Callison-Burch

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 RAG评测 :retrieval augmented generation(title);RAG(abstract);分类 cs.CL

Comments ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07964 2025-06-10 cs.CV cs.AI 79%

SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design

Wenxin Tang, Jingyu Xiao, Wenxuan Jiang, Xi Xiao, Yuhang Wang, Xuxin Tang, Qing Li, Yuehe Ma, Junliang Liu, Shisong Tang, Michael R. Lyu

机构 * Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) Northeastern University(东北大学) Southwest University(西南大学) Kuaishou Technology(快手科技) Peng Cheng Laboratory(鹏城实验室) BNU-HKBU United International College(北京师范大学-香港 Baptist University联合国际学院) Dalian Maritime University(大连海事大学)

专题命中 RAG评测 :RAG(title);retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20871 2025-05-28 cs.CL 79%

Divide-Then-Align: Honest Alignment based on the Knowledge Boundary of RAG

Xin Sun, Jianan Xie, Zhongqi Chen, Qiang Liu, Shu Wu, Yuehe Chen, Bowen Song, Weiqiang Wang, Zilei Wang, Liang Wang

专题命中 RAG评测 :RAG(title,abstract);分类 cs.CL

Comments ACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18225 2025-04-28 cs.CL 79%

Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family

Pierre-Carl Langlais, Pavel Chizhov, Mattia Nee, Carlos Rosas Hinostroza, Matthieu Delsart, Irène Girard, Othman Hicheur, Anastasia Stasenko, Ivan P. Yamshchikov

机构 * PleIAs, Paris, France(巴黎法国PleIAs)

专题命中 RAG评测 :RAG(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18104 2024-10-25 cs.NI cs.AI 79%

ENWAR: A RAG-empowered Multi-Modal LLM Framework for Wireless Environment Perception

Ahmad M. Nazar, Abdulkadir Celik, Mohamed Y. Selim, Asmaa Abdallah, Daji Qiao, Ahmed M. Eltawil

专题命中 RAG评测 :RAG(title);retrieval augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02825 2024-10-15 cs.CL cs.CR 79%

Ingest-And-Ground: Dispelling Hallucinations from Continually-Pretrained LLMs with RAG

Chenhao Fang, Derek Larson, Shitong Zhu, Sophie Zeng, Wendy Summer, Yanqing Peng, Yuriy Hulovatyy, Rajeev Rao, Gabriel Forgues, Arya Pudota, Alex Goncalves, Hervé Robert

专题命中 RAG评测 :RAG(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14501 2026-08-14 cs.SE cs.AI cs.CL 版本更新 79%

CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language

仓颉基准:在低资源通用编程语言上评估大语言模型

Junhang Cheng, Fang Liu, Jia Li, Chengru Wu, Nanxiang Jiang, Li Zhang

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CangjieBench,用于评估大语言模型在低资源通用编程语言上的表现,通过248个高质量样本覆盖文本到代码和代码到代码任务,发现语法受限生成在准确性和计算成本之间取得最佳平衡。

Comments Accepted by ESEM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27654 2026-07-31 cs.CL cs.AI 新提交 79%

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

从单文档到跨文档:大语言模型多粒度事件分析的基准测试

Tao Wen, Shuai Shao, Pei Ke, Xu Han, Jie Zou, Guannan Li, Tao Tian, Jinjie Qiu, Lan Wang, Ke Qin

机构 * University of Electronic Science and Technology of China(电子科技大学) Tsinghua University(清华大学)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文推出MiGUE-Bench基准及MiGUE-Pipeline框架,通过四项核心任务评估LLMs多粒度事件分析能力,明确其能力边界与缺陷,为该领域改进提供方向。

Comments 9 pages. Published in the Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026)

Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), pp. 3464-3472, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03233 2026-07-07 cs.CR cs.AI cs.IR cs.SI 新提交 79%

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

用于开源情报和网络调查的智能与生成式人工智能:分类法、评估、挑战及未来方向

Eduardo Almeida Palmieri, Mohamed Chahine Ghanem, Dipo Dunsin, Zubair Baig, Ed de Quincey, Kim-Kwang Raymond Choo

机构 * School of Computer Science and Mathematics, Keele University(基尔大学计算机科学与数学学院) Cybersecurity Institute, University of Liverpool(利物浦大学网络安全研究所) Department of Applied Computing IICL, University of Wales Trinity Saint David(威尔士特里尼达大学应用计算系) Deakin Cyber Research and Innovation Hub, Deakin University(德金大学网络安全研究与创新中心) Department of Information Systems and Cyber Security, The University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校信息系与网络安全系)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 研究指出公开数字信息增长使人工开源情报分析不足,大语言模型等成解决方案但评估框架滞后。通过综述74项研究,明确智能人工智能类别,找出幻觉验证差距,映射研究到情报生命周期并给出研究议程。

Comments 36 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06142 2026-07-07 cs.CL cs.AI 版本更新 79%

IRC-Bench: Recognizing Entities from Contextual Cues in First-Person Reminiscences

IRC-Bench: 从第一人称回忆中的上下文线索识别实体

Yehudit Aperstein, Eden Moran, Alexander Apartsin

机构 * Intelligent Systems, Afeka Academic College of Engineering(阿法卡学术工程学院智能系统) School of Computer Science, Faculty of Sciences, Holon Institute of Technology(霍隆理工学院计算机科学学院)

专题命中 RAG评测 :RAG(abstract,abstract_cn);dense retrieval(abstract);分类 cs.CL、cs.AI

AI总结 本文提出IRC-Bench基准,用于评估回忆文本中隐式实体识别任务,通过对比本地提及与分散叙述证据,探讨非局部性挑战,测试多种模型配置。

Comments 36 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.04458 2026-06-23 cs.CL cs.IR 版本更新 79%

DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation

DoGMaTiQ:面向报告评估的问答片段自动生成

Bryan Li, William Walden, Yu Hou, Gabrielle Kaili-May Liu, Dawn Lawrie, James Mayfield, Eugene Yang, Chris Callison-Burch, Laura Dietz

机构 * Google Inc.(谷歌公司) Johns Hopkins University(约翰霍普金斯大学) University of Maryland(马里兰大学) Yale University(耶鲁大学) University of Pennsylvania(宾夕法尼亚大学) University of New Hampshire(新罕布什尔大学)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

AI总结 提出DoGMaTiQ流水线,通过文档锚定生成、释义聚类和基于质量准则的子选择三步自动生成高质量QA片段,实现报告的全自动评估,在跨语言任务中与人工判断高度相关。

Comments ICTIR '26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08507 2026-06-16 cs.CL cs.AI 79%

Introducing A Bangla Sentence - Gloss Pair Dataset for Bangla Sign Language Translation and Research

介绍一个孟加拉语句子- gloss配对数据集用于孟加拉语手语翻译和研究

Neelavro Saha, Rafi Shahriyar, Nafis Ashraf Roudra, Saadman Sakib, Annajiat Alim Rasel

机构 * Department of Computer Science and Engineering, Bangladesh University of Engineering and Technology(Bangladesh University of Engineering and Technology计算机科学与工程系)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文介绍了一个包含1000个人工标注句子- gloss配对的新数据集Bangla-SGP,通过规则基于的检索增强生成管道生成约3000个合成配对,用于孟加拉语手语翻译和研究。

Journal ref Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pp. 10457-10466, ELRA, Palma, Mallorca, Spain, May 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13647 2026-06-12 cs.CL cs.AI cs.LG 新提交 79%

SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

SkMTEB:斯洛伐克大规模文本嵌入基准与模型适配

Marek Šuppa, Andrej Ridzik, Daniel Hládek, Natália Kňažeková, Viktória Ondrejová

机构 * Comenius University in Bratislava(布拉迪斯拉发夸美纽斯大学) Cisco Systems(思科系统) Technical University of Košice(科希策技术大学) Kempelen Institute of Intelligent Technologies(肯佩伦智能技术研究所)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 针对低资源西斯拉夫语斯洛伐克语,构建首个MTEB风格文本嵌入基准SkMTEB(含31个数据集、7类任务),并开发高效本地部署模型e5-sk-small/large,通过词汇裁剪与微调在参数减少62%下达到与商业API相当的竞争力。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08873 2026-06-03 cs.IR cs.AI cs.CY cs.SI physics.soc-ph 79%

Whose Name Comes Up? II: Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation

谁的名字出现?II:基于基准测试和干预审计的LLM学者推荐系统

Lisette Espín-Noboa, Gonzalo Gabriel Méndez

机构 * Complexity Science Hub Vienna(维也纳复杂性科学中心) Universitat Politècnica de València(巴塞罗那理工大学) Inria Rennes(里昂国家信息与自动化研究所)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 提出LLMScholarBench基准,通过温度变化、表示约束提示和检索增强生成等干预措施审计22个LLM在物理专家推荐中的技术质量和社会代表性,发现干预措施带来不同权衡。

Comments In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD '26). 30 pages: 11 pages in main (6 figures, 1 table), 19 pages in appendix (22 figures, 2 tables)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31086 2026-06-02 cs.CL cs.IR 79%

Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

超越静态对话:对现实、异构和演化长期记忆的基准测试

Han Zhang, Zihao Tang, Xin Yu, Xiao Liu, Yeyun Gong, Haizhen Huang, Yan Lu, Weiwei Deng, Feng Sun, Qi Zhang, Hanfang Yang

机构 * Center for Applied Statistics, Renmin University of China(中国人民大学应用统计中心) School of Statistics, Renmin University of China(中国人民大学统计学院) Microsoft(微软公司)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

AI总结 针对现有大语言模型记忆基准中对话缺乏长期语义一致性、人物静态化以及忽视异构数据流的问题,提出RHELM基准,通过用户画像和LOOP模块生成具有动态时间演化和长期连贯性的现实对话,并整合异构外部源,评估模型在多源聚合和现实上下文推理方面的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25152 2026-05-27 cs.AI cs.IR 79%

OMD-GraphRAG: Enhancing GraphRAG with Ontology-Guided Extraction, Multi-Dimensional Clustering and Dual-Channel Fusion

OMD-GraphRAG:利用本体引导提取、多维聚类和双通道融合增强GraphRAG

Jie Wang, Honghua Huang, Xi Ge, Jianhui Su, Wen Liu, Shiguo Lian

机构 * Data Science & Artificial Intelligence Research Institute(数据科学与人工智能研究院)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 提出OMD-GraphRAG框架,通过本体引导知识提取、多维社区聚类和双通道图检索融合,提升GraphRAG在复杂推理和多跳查询中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12313 2026-05-13 cs.CL cs.IR 79%

Overview of the MedHopQA track at BioCreative IX: track description, participation and evaluation of systems for multi-hop medical question answering

BioCreative IX MedHopQA 轨道概述:轨道描述、参与及多跳医学问答系统评估

Rezarta Islamaj, Joey Chan, Robert Leaman, Jongmyung Jung, Hyeongsoon Hwang, Quoc-An Nguyen, Hoang-Quynh Le, Harikrishnan Gurushankar Saisudha, Ganesh Chandrasekar, Rustam R. Taktashov, Nadezhda Yu. Bizyukova, Sofia I. R. Conceição, Paulo R. C. Lopes, Reem Abdel Salam, Mary Adewunmi, Zhiyong Lu

机构 * National Library of Medicine (NLM), National Institutes of Health (NIH)(美国国家医学图书馆(NLM)、国家卫生研究院(NIH)) University of Illinois at Urbana Champaign(伊利诺伊大学厄巴纳-香槟分校) Korea University(韩国大学) VNU University of Engineering and Technology, Hanoi, Vietnam(越南河内工程大学) Concordia University, Montreal, QC, CA(蒙特利尔大学) Institute of Biomedical Chemistry (IBMC), 10 bld. 8, Pogodinskaya str., 119121 Moscow, Russia(俄罗斯生物医学化学研究所(IBMC)) LASIGE, Departamento de Informática, Faculdade de Ciências, Universidade de Lisboa, 1749-016 Lisbon, Portugal(葡萄牙里斯本大学 LASIGE 实验室) Faculty of Engineering, Computer Engineering Department Cairo University(埃及开罗大学工程学院) Menzies School of Health Research, Charles Darwin University, NT, Australia(澳大利亚查尔斯达尔文大学梅恩兹健康研究中心) CaresAI, Australia(澳大利亚 CaresAI)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

AI总结 本文介绍了BioCreative IX MedHopQA共享任务,旨在评估大型语言模型在多跳推理中的表现,通过构建1000个挑战性问题,展示了检索增强生成策略的重要性,并提供了公开数据集和评估结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27852 2026-05-01 cs.IR cs.AI 79%

NeocorRAG: Less Irrelevant Information, More Explicit Evidence, and More Effective Recall via Evidence Chains

NeocorRAG:通过证据链减少无关信息,提高明确证据和回忆效果

Shiyao Peng, Qianhe Zheng, Zhuodi Hao, Zichen Tang, Rongjin Li, Qing Huang, Jiayu Huang, Jiacheng Liu, Yifan Zhu, Haihong E

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 NeocorRAG通过证据链系统挖掘和利用,实现检索质量的全面优化,提升召回效果并减少无关信息,取得SOTA性能。

Comments Accepted to WWW 2026

Journal ref Proc. ACM Web Conf. 2026, pages 1899-1910

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24473 2026-04-28 cs.AI cs.CL 79%

Agentic clinical reasoning over longitudinal myeloma records: a retrospective evaluation against expert consensus

基于纵向骨髓瘤记录的代理临床推理:一项回顾性评估与专家共识的对比

Johannes Moll, Jannik Lübberstedt, Christoph Nuernbergk, Jacob Stroh, Luisa Mertens, Anna Purcarea, Christopher Zirn, Zeineb Benchaaben, Fabian Drexel, Hartmut Häntze, Anirudh Narayanan, Friedrich Puttkammer, Andrei Zhukov, Jacqueline Lammert, Sebastian Ziegelmayer, Markus Graf, Marion Högner, Marcus Makowski, Florian Bassermann, Lisa C. Adams, Jiazhen Pan, Daniel Rueckert, Krischan Braitsch, Keno K. Bressem

机构 * Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM) and TUM University Hospital(人工智能在医疗与健康中的研究所,慕尼黑技术大学(TUM)及慕尼黑技术大学医院) Department of Diagnostic and Interventional Radiology, Klinikum rechts der Isar, TUM University Hospital, School of Medicine and Health, Technical University of Munich(诊断与介入放射科,莱茵河右岸医院,慕尼黑技术大学医院,医学与健康学院,慕尼黑技术大学) Department of Cardiovascular Radiology and Nuclear Medicine, German Heart Center, TUM University Hospital, School of Medicine and Health, Technical University of Munich(心血管放射学与核医学科,德国心脏中心,慕尼黑技术大学医院,医学与健康学院,慕尼黑技术大学) Department of Medicine III, Klinikum rechts der Isar, TUM University Hospital, School of Medicine and Health, Technical University of Munich(第三医学部,莱茵河右岸医院,慕尼黑技术大学医院,医学与健康学院,慕尼黑技术大学)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本研究评估了代理推理系统在骨髓瘤长期记录中的表现,发现其在复杂问题和长记录中优于传统方法,但需进一步验证以确保临床应用的安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16518 2026-04-28 cs.CL cs.AI 79%

CUB: Benchmarking Context Utilisation Techniques for Language Models

CUB:语言模型上下文利用技术的基准测试

Lovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee, Richard Johansson, Hyunsoo Cho, Isabelle Augenstein

机构 * Chalmers University of Technology(查尔姆斯理工大学) University of Gothenburg(哥德堡大学) Seoul National University(首尔国立大学) University of Copenhagen(哥本哈根大学) Ewha Womans University(成均馆大学)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CUB基准,用于评估语言模型在不同噪声上下文下的上下文利用技术,发现现有方法在真实场景中存在不足,需更全面的测试。

Comments Accepted at ACL 2026, 33 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20763 2026-04-23 cs.IR cs.AI cs.LG 79%

Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation

覆盖,而非平均:语义分层用于可信检索评估

Andrew Klearman, Radu Revutchi, Rohin Garg, Rishav Chakravarti, Samuel Marc Denton, Yuan Xue

机构 * Scale AI

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文提出语义分层方法,通过构建可解释的实体聚类空间,解决检索评估中覆盖性和稳定性问题,提升评估的透明度和可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05527 2025-11-11 cs.CL cs.AI 79%

Bridging Industrial Expertise and XR with LLM-Powered Conversational Agents

Despina Tomkou, George Fatouros, Andreas Andreou, Georgios Makridis, Fotis Liarokapis, Dimitrios Dardanis, Athanasios Kiourtis, John Soldatos, Dimosthenis Kyriazis

机构 * Innov-Acts Ltd.(Innov-Acts有限公司) CYENS Centre of Excellence(CYENS卓越中心) University of Piraeus(比雷埃克斯大学)

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract);knowledge retrieval(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 7 figures

Journal ref 2025 21st International Conference on Distributed Computing in Smart Systems and the Internet of Things (DCOSS-IoT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04468 2025-09-08 cs.CL cs.AI 79%

Evaluating Large Language Models for Financial Reasoning: A CFA-Based Benchmark Study

Xuan Yao, Qianteng Wang, Xinbo Liu, Ke-Wei Huang

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract);knowledge retrieval(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14063 2025-08-21 cs.IR cs.AI 79%

A Multi-Agent Approach to Neurological Clinical Reasoning

Moran Sorka, Alon Gorenshtein, Dvir Aran, Shahar Shelly

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract);knowledge retrieval(abstract);分类 cs.IR、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10586 2025-07-16 cs.CL cs.AI 79%

AutoRAG-LoRA: Hallucination-Triggered Knowledge Retuning via Lightweight Adapters

Kaushik Dwivedi, Padmanabh Patanjali Mishra

机构 * BITS Pilani(比斯·皮兰大学) University of Adelaide(阿德莱德大学)

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract);hybrid retrieval(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.21033 2025-07-16 cs.CL cs.AI 79%

Plancraft: an evaluation dataset for planning with LLM agents

Gautier Dagan, Frank Keller, Alex Lascarides

机构 * University of Edinburgh(爱丁堡大学)

专题命中 RAG评测 :retrieval augmented generation(abstract);RAG(abstract);retriever(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏