arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 1199 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 1199 篇

2412.05206 2024-12-09 cs.CL cs.AI cs.IR 67%

ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges

Kaustubh D. Dhole, Kai Shu, Eugene Agichtein

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00495 2024-12-03 cs.GT cs.LG 67%

Rethinking Strategic Mechanism Design In The Age Of Large Language Models: New Directions For Communication Systems

Ismail Lotfi, Nouf Alabbasi, Omar Alhussein

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract)

Comments submitted to IEEE IoTM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14926 2024-10-22 cs.CE 67%

Aligning LLMs with Human Instructions and Stock Market Feedback in Financial Sentiment Analysis

Zijie Zhao, Roy E. Welsch

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15232 2024-10-21 cs.CL cs.AI cs.IR 67%

Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations

Yucheng Jiang, Yijia Shao, Dekun Ma, Sina J. Semnani, Monica S. Lam

专题命中 RAG评测 :RAG(abstract);分类 cs.IR、cs.CL、cs.AI

Comments EMNLP 2024 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16605 2024-09-26 cs.CL cs.AI cs.IR cs.LG 67%

Evaluating and Enhancing Large Language Models for Novelty Assessment in Scholarly Publications

Ethan Lin, Zhiyuan Peng, Yi Fang

专题命中 RAG评测 :RAG(abstract);分类 cs.IR、cs.CL、cs.AI

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07487 2024-09-17 q-fin.CP 67%

MoA is All You Need: Building LLM Research Team using Mixture of Agents

Sandy Chen, Leqi Zeng, Abhinav Raghunathan, Flora Huang, Terrence C. Kim

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01701 2024-06-21 cs.SE 67%

De-Hallucinator: Mitigating LLM Hallucinations in Code Generation Tasks via Iterative Grounding

Aryaz Eghbali, Michael Pradel

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10018 2024-06-17 cs.SE 67%

STALL+: Boosting LLM-based Repository-level Code Completion with Static Analysis

Junwei Liu, Yixuan Chen, Mingwei Liu, Xin Peng, Yiling Lou

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract)

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01768 2024-06-07 cs.NI cs.IT eess.SP math.IT 67%

TSpec-LLM: An Open-source Dataset for LLM Understanding of 3GPP Specifications

Rasoul Nikbakht, Mohamed Benzaghta, Giovanni Geraci

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16543 2024-05-22 cs.AR 67%

RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language Models

Yun-Da Tsai, Mingjie Liu, Haoxing Ren

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04662 2024-05-21 cs.CR 67%

VulLibGen: Generating Names of Vulnerability-Affected Packages via a Large Language Model

Tianyu Chen, Lin Li, Liuchuan Zhu, Zongyang Li, Xueqing Liu, Guangtai Liang, Qianxiang Wang, Tao Xie

专题命中 RAG评测 :retrieval augmented generation(abstract);RAG(abstract)

Comments ACL 2024 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12430 2024-02-27 cs.IR cs.AI cs.CL cs.LG 67%

Efficient Title Reranker for Fast and Improved Knowledge-Intense NLP

Ziyi Chen, Jize Jiang, Daqian Zuo, Heyi Tao, Jun Yang, Yuxiang Wei

专题命中 RAG评测 :RAG(abstract);分类 cs.IR、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14480 2024-02-23 cs.SE 67%

MeTMaP: Metamorphic Testing for Detecting False Vector Matching Problems in LLM Augmented Generation

Guanyu Wang, Yuekang Li, Yi Liu, Gelei Deng, Tianlin Li, Guosheng Xu, Yang Liu, Haoyu Wang, Kailong Wang

专题命中 RAG评测 :retrieval-augmented generation(abstract);RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17041 2026-07-02 cs.CL cs.IR 新提交 66%

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio

对Nature Portfolio元分析文章进行LLM代理基准测试

Anzhe Xie, Weihang Su, Yujia Zhou, Yiqun Liu, Min Zhang, Qingyao Ai

机构 * Tsinghua University(清华大学)

专题命中 RAG评测 :RAG(abstract_cn);分类 cs.IR、cs.CL;retriever(comments)

AI总结 提出MetaSyn数据集,包含442篇专家策划的元分析,用于评估LLM代理在检索-筛选-综合全流程中的表现,发现当前系统在筛选阶段存在严重瓶颈。

Comments 17 pages, 8 figures, 11 tables. Code and evaluation: https://github.com/THUIR/MetaSyn Dataset: https://huggingface.co/datasets/THUIR/MetaSyn Model: https://huggingface.co/BFTree/MA-Retriever

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00009 2026-08-04 cs.CL cs.AI 新提交 62%

AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents

AgentMemBench:用于评估对话AI智能体长期记忆管理策略的系统基准

Ahmed Cherif

机构 * Sofrecom(索弗雷科姆)

专题命中 RAG评测 :dense retrieval(abstract);分类 cs.CL、cs.AI

AI总结 该研究构建了AgentMemBench基准,评估五种记忆策略及两款现有系统,发现外部键值存储(EKV)在长程对话记忆任务中表现最优,但存在内存占用成本,同时公开了全部可复现资源。

Comments 22 pages, 3 figures submitted on Neural Computing and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24802 2026-07-31 cs.IR cs.CL 版本更新 62%

SourceMinds at CheckThat! 2026: NLI-Grounded Citation Auditing in a Multi-Agent Pipeline for Full Fact-Checking Article Generation

SourceMinds参与2026年CheckThat!:多智能体管道中基于自然语言推理的引用审核用于完整事实核查文章生成

Farhan Sharukh Hasan, Anirban Saha Anik, Eric Liu, Xiaoying Song, Mohotarema Rashid, Lingzi Hong

专题命中 RAG评测 :dense retrieval(abstract);分类 cs.IR、cs.CL

AI总结 针对CLEF 2026 CheckThat!实验室任务3,提出多智能体管道系统,结合证据检索、结构化规划、文章生成、自我批判及引用审核等方法,强调证据选择、结构化生成与生成后引用验证对事实核查文章生成的重要性。

Comments CLEF 2026 Working Notes / CheckThat! Lab at CLEF 2026, Jena, Germany

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07314 2026-07-30 cs.CL cs.AI 版本更新 62%

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

MEDIC:对LLM在临床应用中的安全性和实用性领先指标的综合评估

Praveenkumar Kanithi, Clément Christophe, Marco AF Pimentel, Tathagata Raha, Prateek Munjal, Nada Saadi, Hamza A Javed, Svetlana Maslenkova, Nasir Hayat, Ronnie Rajan, Shadab Khan

机构 * M42

专题命中 RAG评测 :knowledge retrieval(abstract);分类 cs.CL、cs.AI

AI总结 MEDIC通过综合评估框架揭示LLM在临床应用中的安全性和实用性差异,强调需采用组合方法以应对多维度性能权衡。

Comments Published in Transactions on Machine Learning Research (06/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24651 2026-07-28 cs.CV cs.CL cs.IR 新提交 62%

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

无坐标或区域标签的视觉文档理解中的证据归因

Zhuchenyang Liu, Yao Zhang, Yu Xiao

机构 * Aalto University(阿尔托大学)

专题命中 RAG评测 :retriever(abstract);分类 cs.IR、cs.CL

AI总结 研究视觉文档理解中无坐标或区域标签时的证据归因问题,通过对比坐标与语言接口,发现语言接口能提升证据召回率、降低幻觉率。基于此用引述和检索管道作训练框架,引入GRPO方法,提高了模型严格归因准确率,找到改善归因的实用路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21274 2026-07-24 cs.CL cs.AI 新提交 62%

A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset

希腊图书出版商环境中嵌入模型和大语言模型的比较评估 - CUP数据集

Katerina Papantoniou, Panagiotis Papadakos, Theodore Patkos, Dimitris Garefalakis, Nikos Vardakis, Dimitris Plexousakis

机构 * ICS-FORTH(希腊计算机科学与技术研究所) Crete University Press(克里特大学出版社)

专题命中 RAG评测 :hybrid retrieval(abstract);分类 cs.CL、cs.AI

AI总结 该研究提出希腊图书检索基准CUP数据集,评估稀疏、密集、混合及大语言模型辅助的检索方法,发现多语言嵌入优于希腊特定模型,混合检索最佳,还分析了不同方法在各类查询中的表现及大语言模型相关技术的效果。

Comments Preprint of a manuscript submitted to the 14th EETN Conference on Artificial Intelligence (SETN 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20431 2026-07-24 cs.CL cs.IR 新提交 62%

Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis

用于证据感知材料文献分析的技能收缩代理

Bixuan Li, Yu Liu, Shuo Shi, Xiaoya Huang, Peng Kang, Lei Zheng

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

AI总结 研究针对材料科学文献分析难题,提出技能驱动的AlphaAgent框架,通过明确技能契约解耦任务。其含专门检索技能和报告生成技能,在盲评中显著优于基线系统,提升了文献分析效果。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17288 2026-07-21 cs.DB cs.AI 新提交 62%

SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation

SAGA:用于时间基准生成的合成智能体图架构

Jiacheng Ding, Xiaofei Zhang

机构 * University of Memphis(密苏里大学)

专题命中 RAG评测 :RAG(abstract);分类 cs.AI、cs.DB

AI总结 针对时间图基准稀缺问题,SAGA系统通过四阶段管道生成大规模语义丰富的时间图,其架构解耦结构与语义,能实现结构真实性、语义丰富性和自动异常标注,在单个H100 GPU上高效生成大量带受控异常的时间边。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01086 2026-07-17 cs.AI cs.CR cs.DB cs.DC cs.SE 版本更新 62%

MedBeads: An AI-Native Clinical Context Graph Built from Immutable Beads and Reconstructable Clinical Links

MedBeads:面向可信医疗AI的智能体原生不可变数据基底

Takahito Nakajima

机构 * Diagnostic Imaging and Interventional Radiology, Institute of Medicine, University of Tsukuba(东京大学医学研究院诊断影像与介入放射学部) Center for Cyber Medicine Research, University of Tsukuba(东京大学计算机医学研究中心)

专题命中 RAG评测 :RAG(abstract_cn);分类 cs.AI、cs.DB

AI总结 针对医疗AI中电子病历与智能体间的上下文不匹配问题,提出基于Merkle有向无环图的不可变数据架构MedBeads,通过确定性图遍历替代概率检索,实现可审计、防篡改的临床上下文提供。

Comments 23 pages, 5 figures, 3 tables. Reference implementation and reproducible Docker demo available at https://github.com/medbeads/medbeads

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00761 2026-07-17 cs.AI cs.CL 版本更新 62%

L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit

L-MARS:具有协调推理和代理搜索的法律多智能体工作流

Boqin Yuan, Ziqi Wang

机构 * University of Southern California(南加州大学) University of California, San Diego(加利福尼亚大学圣迭戈分校)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

AI总结 L-MARS是一种多代理检索框架,通过分解查询为结构化子问题,利用代理网络搜索获取证据,并通过验证代理过滤结果,从而在法律问答中实现高准确率。

Comments Accepted at the AI4Law Workshop at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05456 2026-07-08 cs.AI cs.CL q-bio.QM 新提交 62%

Prompt-to-Paper: Agentic AI System for Bioinformatics

从提示到论文:用于生物信息学的智能AI系统

Ramsha Kamran, Maheera Amjad, Zartasha Mustansar, Arsalan Shaukat, Salma Sherbaz, Muhammad U. S. Khan

机构 * School of Interdisciplinary Engineering and Sciences (SINES), National University of Sciences and Technology (NUST)(跨学科工程与科学学院(SINES),国立科技大学(NUST))

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 研究针对现有AI自动生成稿件系统的缺陷,提出Prompt-to-Paper多智能体框架。通过检索增强生成、自主编码实验及八维质量评分等创新,经质量驱动改进循环,在生物信息学案例中验证,有效提升稿件质量,成本低且获较好外部评价。

Comments NA

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28510 2026-06-09 cs.SE cs.AI cs.IR 版本更新 62%

Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets

高效可扩展的LLM生成代码片段溯源追踪

Andrea Gurioli, Davide D'Ascenzo, Federico Pennino, Maurizio Gabbrielli, Stefano Zacchiroli

机构 * University of Bologna(博洛尼亚大学)

专题命中 RAG评测 :vector search(abstract);分类 cs.IR、cs.AI

AI总结 提出混合两阶段溯源追踪流水线HYBRIDSOURCETRACKER,结合向量搜索与指纹匹配,实现LLM生成代码的高效、可扩展溯源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04704 2026-05-29 cond-mat.mtrl-sci cs.AI cs.CL 62%

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials

AtomWorld: 评估大型语言模型在晶体材料空间推理能力的基准

Taoyuze Lv, Alexander Chen, Fengyu Xie, Chu Wu, Jeffrey Meng, Dongzhan Zhou, Yingheng Wang, Bram Hoex, Zhicheng Zhong, Tong Xie

机构 * University of New South Wales, NSW, Sydney, Australia(新南威尔士大学,新州,悉尼,澳大利亚) Suzhou Institute for Advanced Research, University of Science(苏州先进研究院,科学大学) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国) Cornell University(康奈尔大学)

专题命中 RAG评测 :knowledge retrieval(abstract);分类 cs.CL、cs.AI

AI总结 提出AtomWorld基准,通过十种基本原子结构操作评估LLM在材料科学中的空间推理能力,发现Claude Opus 4.6表现最佳但复杂空间关系操作成功率低,表明LLM更适合作为辅助工具而非完全自主的科研代理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19341 2026-05-20 cs.CL cs.AI cs.LG stat.ML 62%

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models

HalluWorld: 一个用于通过参考世界模型控制幻觉的基准

Emmy Liu, Varun Gangal, Michael Yu, Zhuofu Tao, Karan Singh, Sachin Kumar, Steven Y. Feng

机构 * Carnegie Mellon University(卡内基梅隆大学) Patronus AI Independent Researcher(独立研究者) Stanford University(斯坦福大学) The Ohio State University(俄亥俄州立大学) DegenAI Labs(DegenAI实验室)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出HalluWorld基准,通过显式参考世界模型研究语言模型的幻觉问题,发现不同任务中幻觉表现不一致,表明幻觉源于多种失败模式而非单一能力。

Comments HalluWorld benchmark (code and data) at github.com/DegenAI-Labs/HalluWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13295 2026-05-14 cs.CL cs.AI cs.MA 62%

CANTANTE: Optimizing Agentic Systems via Contrastive Credit Attribution

CANTANTE:通过对比信用分配优化代理系统

Tom Zehle

机构 * University of Freiburg(弗赖堡大学) ELLIS Institute(埃里克·林斯研究所) Tübingen(图宾根)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CANTANTE框架,通过对比多个联合配置的rollouts分解系统奖励为单个代理的更新信号,提升多代理系统优化效果,在编程、数学推理和多跳问答任务中均取得最佳排名。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14952 2026-04-28 cs.CL cs.AI 62%

CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning

CorpusQA:一个1000万词的语料库分析与推理基准

Zhiyuan Lu, Chenliang Li, Yingcheng Shi, Weizhou Shen, Ming Yan, Fei Huang

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 CorpusQA提出一个1000万词的基准,解决大规模语料库推理问题,通过合成数据框架生成复杂查询,验证长文本推理能力,发现先进架构对全局信息整合的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20869 2026-04-24 cs.CY cs.AI cs.HC cs.IR cs.LG 62%

Clinical Reasoning AI for Oncology Treatment Planning: A Multi-Specialty Case-Based Evaluation

肿瘤治疗规划的临床推理AI:多专科基于病例的评估

Philippe E. Spiess, Md Muntasir Zitu, Alison Walker, Daniel A. Anaya, Robert M. Wenham, Michael Vogelbaum, Daniel Grass, Ali-Musa Jaffer, Amod Sarnaik, Caitlin McMullen, Christine Sam, John V. Kiluk, Tianshi Liu, Tiago Biachi, Julio Powsang, Jing-Yi Chern, Roger Li, Seth Felder, Samuel Reynolds, Michael Shafique, Alison Sheehan, Ashley Layman, Cydney A. Warfield, Derrick Legoas, Jaclyn Parrinello, Jena Schmitz, Kevin Eaton, Mark Honor, Luis Felipe, Issam ElNaqa, Elier Delgado, Talia Berler, Rachael V. Phillips, Frantz Francisque, Carlos Garcia Fernandez, Gilmer Valdes

机构 * Moffitt Cancer Center and Research Institute(莫菲特癌症中心与研究学院) Innova Montréal Inc.(蒙特利尔创新公司) Oncobrain, Inc.(Oncobrain公司) Advanced Cancer Treatment Centers(先进癌症治疗中心)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文评估了OncoBrain在多专科病例中的表现,展示了其在肿瘤治疗规划中的准确性、安全性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏