OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL
AI 大模型
检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI
Comments preprint
专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL
Comments Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), 2025
专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL
Comments 23pages
专题命中 RAG评测 :retrieval augmented generation(title);retrieval-augmented generation(abstract);RAG(abstract);分类 cs.IR
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL
Comments NeurIPS 2024 Datasets and Benchmarks Track
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL
Comments Findings of EMNLP Camera-ready version
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR
专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.CL
Comments 11 pages, 5 figures, 9 tables
专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.AI
专题命中 RAG评测 :RAG(title,abstract);retriever(abstract);分类 cs.CL
专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.AI
专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.CL
专题命中 RAG评测 :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.CL
Comments These authors contributed equally to this work
通过语言强化学习实现大语言模型个性化的偏好适配学习
专题命中 RAG评测 :RAG(summary_cn,abstract);分类 cs.CL、cs.AI
AI总结 本研究针对LLM个性化中通用偏好摘要冗余问题,提出无训练元学习框架AlignXada,经语言强化学习优化后在多任务多模型上提升性能且适配效果优于RAG。
IDP AutoOpt:智能文档处理流水线配置的智能体驱动优化
专题命中 RAG评测 :RAG(summary_cn,abstract);分类 cs.IR、cs.AI
AI总结 IDP AutoOpt是自主LLM智能体,通过闭环流程优化IDP流水线配置,在多领域任务上性能优于人类专家且成本更低,还可扩展至RAG等其他企业AI系统。
LLM 能时间旅行吗?通过强化学习增强法律智能搜索中的时间一致性
机构 * Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China(香港科技大学计算机科学与工程系) ; School of Law, Tsinghua University, Beijing, China(清华大学法学院) ; Cheriton School of Computer Science, University of Waterloo, Waterloo, Canada(滑铁卢大学丘成桐计算机科学系)
专题命中 RAG评测 :RAG(summary_cn,abstract);分类 cs.CL、cs.AI
AI总结 提出 LegalSearch-R1 框架,结合本地 statute RAG 和在线搜索,通过强化学习在跨修订期数据上训练,以解决法律 LLM 的时间偏差和搜索代理缺乏时间约束的问题,在13项法律任务上超越现有方法。
Comments Under Review
上下文感知解码减少查询导向摘要中的幻觉
机构 * University of Utah(犹他大学)
专题命中 RAG评测 :retrieval augmented generation(abstract);RAG(abstract);retriever(abstract);dense retrieval(abstract)
AI总结 本文提出上下文感知解码方法,通过减少查询导向摘要中的幻觉并保留词法模式匹配度,提升生成质量。
Comments technical report
主动检索生成式人工智能何时应进行检索?效用、校准和成本的预算感知评估
机构 * Carnegie Mellon University(卡内基梅隆大学) ; University of Glasgow(格拉斯哥大学) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 RAG评测 :RAG(title,abstract)
AI总结 研究主动检索生成式人工智能何时检索,通过将其重述为效用估计进行预算感知评估,并分离出相关三个问题,利用多种方法实现,在多数据集和模型中验证,强调评估应报告多方面指标。
Comments Accepted at the ACM SIGKDD KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI; 7 pages, 1 figure, and 4 tables
EmbC-Test: 如何利用LLMs和RAG加速嵌入式软件测试
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract)
AI总结 本文提出利用LLMs和RAG技术,通过生成自动测试用例,显著提高嵌入式软件测试效率,节省66%的测试时间并每小时生成270个测试用例。
Journal ref Technical University of Munich. 2026. ISBN 978-3-911430-12-8. https://mediatum.ub.tum.de/1846559
SOSecure: 借助检索增强生成与StackOverflow讨论实现更安全的代码生成
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract)
AI总结 SOSecure通过检索增强生成与StackOverflow讨论提升代码安全性,实现71.7%-96.7%的修复率,优于其他基线方法。
隐于 plain-text:一种用于 RAG 中社会网络间接提示注入的基准测试
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract)
AI总结 OpenRAG-Soc 提供了一种用于评估 RAG 系统在社会网络间接提示注入攻击下的基准测试工具,通过标准化的端到端评估和可部署的缓解措施,帮助从业者跟踪风险并增强部署安全性。
Comments WWW 2026
多阶段验证导向框架用于缓解多模态RAG中的幻觉
机构 * The University of New South Wales(新南威尔士大学)
专题命中 RAG评测 :RAG(title,abstract);分类 cs.IR、cs.CL、cs.AI
AI总结 本文提出了一种多阶段验证导向框架,通过优先考虑事实准确性和真实性来缓解多模态RAG中的幻觉问题,并在KDD Cup 2025中取得第三名。
Comments KDD Cup 2025 Meta CRAG-MM Challenge: Third Prize in the Single-Source Augmentation Task
在RAG系统中减少能耗的所提技术有效性:一项受控实验
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract)
AI总结 本研究通过受控实验评估了五种减少RAG系统能耗的技术,发现调整检索阈值和减少嵌入尺寸能有效降低能耗与延迟,同时保持准确率。
Comments Accepted for publication at the 2026 International Conference on Software Engineering: Software Engineering in Society (ICSE-SEIS'26)
CyberLLM-FINDS 2025:基于检索增强生成和图集成的领域特定LLM指令微调方法用于MITRE评估
专题命中 RAG评测 :retrieval-augmented generation(title,abstract);RAG(abstract)
AI总结 本文提出了一种基于检索增强生成和图集成的领域特定LLM微调方法,通过STIX威胁情报实现与MITRE ATT&CK技术的对齐,提升网络安全威胁情报分析的准确性。
Comments 12 pages
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract)
机构 * Meta Reality Labs(Meta现实实验室) ; Meta Superintelligence Labs(Meta超智能实验室) ; FAIR, Meta(FAIR,Meta) ; Meta
专题命中 RAG评测 :RAG(title,abstract);retrieval-augmented generation(abstract)
机构 * Machine Learning, ICML(机器学习,ICML)
专题命中 RAG评测 :RAG(title);retrieval-augmented generation(abstract);hybrid retrieval(abstract)