arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 1205 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. RAG评测 1205 篇

2510.10390 2025-10-14 cs.CL cs.AI cs.LG 62%

RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models

Aashiq Muhamed, Leonardo F. R. Ribeiro, Markus Dreyer, Virginia Smith, Mona T. Diab

机构 * Carnegie Mellon University(卡内基梅隆大学) Amazon AGI(亚马逊人工智能研究院)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01259 2025-10-03 cs.CL cs.AI cs.CR 62%

In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b

Nils Durner

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

Comments 27 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01237 2025-10-03 cs.CL cs.AI 62%

Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation

Nandakishor M

机构 * AI Safety Research(人工智能安全研究) Convai Innovations(Convai创新)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09598 2025-09-30 cs.CL cs.AI cs.CY 62%

How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation

Ruohao Guo, Wei Xu, Alan Ritter

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22658 2025-09-30 cs.IR cs.AI 62%

How good are LLMs at Retrieving Documents in a Specific Domain?

Nafis Tanveer Islam, Zhiming Zhao

机构 * MultiScale Networked Systems (MNS) University of Amsterdam, Netherlands, University of Amsterdam, Netherlands(多尺度网络化系统(MNS)阿姆斯特丹大学,荷兰,阿姆斯特丹大学,荷兰)

专题命中 RAG评测 :RAG(abstract);分类 cs.IR、cs.AI

Comments Accepted at FAIEMA Conference 2025. DOI will be provided once the conference publishes the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13144 2025-09-29 cs.CL cs.AI 62%

DialSim: A Dialogue Simulator for Evaluating Long-Term Multi-Party Dialogue Understanding of Conversational Agents

Jiho Kim, Woosog Chay, Hyeonji Hwang, Daeun Kyung, Hyunseung Chung, Eunbyeol Cho, Yeonsu Kwon, Yohan Jo, Edward Choi

机构 * KAIST(韩国科学技术院) SNU(釜山大学)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11507 2025-09-26 cs.AI cs.CL 62%

TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains

Wanying Wang, Zeyu Ma, Xuhong Wang, Yangchun Zhang, Pengfei Liu, Mingang Chen

机构 * Shanghai Key Laboratory of Computer Software Testing and Evaluating, Shanghai Development Center of Computer Software Technology(上海软件测试与评估关键实验室、上海计算机软件技术发展中心) Shanghai Normal University(上海师范大学) Shanghai Artificial Intelligence Lab(上海人工智能实验室) Shanghai University(上海大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

Comments Wang et al. Copyright 2026 lEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, including reprinting/republishing, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work. DOI will be added upon IEEE Xplore publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12621 2025-09-25 cs.CL cs.IR 62%

SAFE: Improving LLM Systems using Sentence-Level In-generation Attribution

João Eduardo Batista, Emil Vatai, Mohamed Wahib

机构 * RIKEN-CCS Kobe, Japan(日本神户RIKEN-CCS)

专题命中 RAG评测 :RAG(abstract);分类 cs.IR、cs.CL

Comments 30 pages (9 pages of content, 5 pages of references, 16 pages of supplementary material), 7 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04410 2025-09-23 cs.AI cond-mat.mtrl-sci cs.CL 62%

Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science

Peter Jansen, Samiah Hassan, Ruoyao Wang

机构 * University of Arizona(亚利桑那大学) Allen Institute for Artificial Intelligence(人工智能 Allen 机构)

专题命中 RAG评测 :retrieval augmented generation(abstract);分类 cs.CL、cs.AI

Comments 9 pages (Accepted to EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09848 2025-08-15 cs.CL cs.AI 62%

PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts

Mo Yu, Tsz Ting Chung, Chulun Zhou, Tong Li, Rui Lu, Jiangnan Li, Liyan Xu, Haoshu Lu, Ning Zhang, Jing Li, Jie Zhou

机构 * WeChat AI, Tencent(腾讯微信AI)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

Comments First 7 authors contributed equally. Project page: https://gorov.github.io/prelude

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08942 2025-08-13 cs.CL cs.IR 62%

Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens

Lucas Albarede, Jose Moreno, Lynda Tamine, Luce Lefeuvre

机构 * Dir. Technologies Innovation, SNCF(技术创新部,法国国家铁路公司)

专题命中 RAG评测 :RAG(abstract);分类 cs.IR、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06729 2025-08-12 cs.CL cs.AI 62%

Large Language Models for Oral History Understanding with Text Classification and Sentiment Analysis

Komala Subramanyam Cherukuri, Pranav Abishai Moses, Aisa Sakata, Jiangping Chen, Haihua Chen

机构 * University of North Texas(北卡罗来纳州立大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06600 2025-08-12 cs.CL cs.IR 62%

BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent

Zijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie, Kai Zou, Andrew Liu, Joshua Green, Kshama Patel, Ruoxi Meng, Mingyi Su, Sahel Sharifymoghaddam, Yanxi Li, Haoran Hong, Xinyu Shi, Xuye Liu, Nandan Thakur, Crystina Zhang, Luyu Gao, Wenhu Chen, Jimmy Lin

机构 * University of Waterloo(滑铁卢大学) CSIRO(澳大利亚联邦科学与工业研究组织) Independent(独立研究者) Carnegie Mellon University(卡内基梅隆大学) The University of Queensland(昆士兰大学)

专题命中 RAG评测 :retriever(abstract);分类 cs.IR、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22915 2025-08-01 cs.CL cs.AI 62%

Theoretical Foundations and Mitigation of Hallucination in Large Language Models

Esmail Gumaan

机构 * Department of Computer Science, University of Sana'a(桑加大学计算机科学系)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14651 2025-07-29 cs.CL cs.AI 62%

Real-time Factuality Assessment from Adversarial Feedback

Sanxing Chen, Yukun Huang, Bhuwan Dhingra

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18392 2025-07-25 cs.CL cs.AI cs.LG 62%

CLEAR: Error Analysis via LLM-as-a-Judge Made Easy

Asaf Yehudai, Lilach Eden, Yotam Perlitz, Roy Bar-Haim, Michal Shmueli-Scheuer

机构 * IBM Research(IBM研究院) The Hebrew University of Jerusalem(特拉维夫大学)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22486 2025-07-01 cs.CL cs.AI 62%

Hallucination Detection with Small Language Models

Ming Cheung

机构 * dBeta Labs, The Lane Crawford Joyce Group(dBeta实验室,Lane Crawford Joyce集团)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

Journal ref Hallucination Detection with Small Language Models, IEEE International Conference on Data Engineering (ICDE), Workshop, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01344 2025-06-30 cs.AI cs.CL 62%

Pairing Analogy-Augmented Generation with Procedural Memory for Procedural Q&A

K Roth, Rushil Gupta, Simon Halle, Bang Liu

机构 * Université de Montréal & Mila(蒙特利尔大学及Mila) Thales Canada(加拿大泰雷兹公司) Canada CIFAR AI Chair(加拿大CIFAR人工智能 chair)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19175 2025-06-16 cs.CL cs.AI 62%

MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential Diagnosis

Daniel Rose, Chia-Chien Hung, Marco Lepri, Israa Alqassem, Kiril Gashteovski, Carolin Lawrence

专题命中 RAG评测 :knowledge retrieval(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04388 2025-05-30 cs.CL cs.AI 62%

The Aloe Family Recipe for Open and Specialized Healthcare LLMs

Dario Garcia-Gasulla, Jordi Bayarri-Planas, Ashwin Kumar Gururajan, Enrique Lopez-Cuena, Adrian Tormos, Daniel Hinjos, Pablo Bernabeu-Perez, Anna Arias-Duart, Pablo Agustin Martin-Torres, Marta Gonzalez-Mallo, Sergio Alvarez-Napagao, Eduard Ayguadé-Parra, Ulises Cortés

机构 * Barcelona Supercomputing Center (BSC-CNS)(巴塞罗那超级计算中心(BSC-CNS)) Universitat Politècnica de Catalunya - Barcelona Tech (UPC)(加泰罗尼亚理工大学-巴塞罗那技术大学(UPC))

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

Comments Follow-up work from arXiv:2405.01886

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19494 2025-05-27 cs.CL cs.IR 62%

Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents

Manoj Balaji Jagadeeshan, Prince Raj, Pawan Goyal

机构 * Indian Institute of Technology, Kharagpur(印度理工学院,克拉格浦尔)

专题命中 RAG评测 :RAG(abstract);分类 cs.IR、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11571 2025-05-20 cs.CL cs.IR cs.LG 62%

FaMTEB: Massive Text Embedding Benchmark in Persian Language

Erfan Zinvandi, Morteza Alikhani, Mehran Sarmadi, Zahra Pourbahman, Sepehr Arvin, Reza Kazemi, Arash Amini

机构 * MCINext Sharif University of Technology(谢里夫理工大学)

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21716 2025-05-01 cs.RO cs.AI cs.CL 62%

LLM-Empowered Embodied Agent for Memory-Augmented Task Planning in Household Robotics

Marc Glocker, Peter Hönig, Matthias Hirschmanner, Markus Vincze

机构 * Automation and Control Institute, Faculty of Electrical Engineering, TU Wien(自动化与控制研究所,电气工程学院,维也纳技术大学) AIT Austrian Institute of Technology GmbH, Center for Vision, Automation and Control(奥地利技术研究所(AIT)有限公司,视觉、自动化与控制中心)

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

Comments Accepted at Austrian Robotics Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18110 2025-04-23 cs.CL cs.IR 62%

Open-World Evaluation for Retrieving Diverse Perspectives

Hung-Ting Chen, Eunsol Choi

机构 * Department of Computer Science(计算机科学系) New York University(纽约大学)

专题命中 RAG评测 :retriever(abstract);分类 cs.IR、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12170 2025-04-21 cs.CL cs.AI 62%

Where is the answer? Investigating Positional Bias in Language Model Knowledge Extraction

Kuniaki Saito, Kihyuk Sohn, Chen-Yu Lee, Yoshitaka Ushiku

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

Comments Code is published at https://github.com/omron-sinicx/WhereIsTheAnswer

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16063 2025-04-18 cs.CL cs.AI 62%

Citation-Enhanced Generation for LLM-based Chatbots

Weitao Li, Junkai Li, Weizhi Ma, Yang Liu

专题命中 RAG评测 :retrieval augmented generation(abstract);分类 cs.CL、cs.AI

Journal ref Proc. 62nd ACL Vol. 1 Long Papers, Bangkok Thailand, pp. 1451-1466, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10179 2025-04-15 cs.AI cs.CL cs.ET 62%

The Future of MLLM Prompting is Adaptive: A Comprehensive Experimental Evaluation of Prompt Engineering Methods for Robust Multimodal Performance

Anwesha Mohanty, Venkatesh Balavadhani Parthasarathy, Arsalan Shahid

专题命中 RAG评测 :knowledge retrieval(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11989 2025-03-21 cs.CL cs.AI 62%

Applications of Large Language Model Reasoning in Feature Generation

Dharani Chandra

专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

Comments I just updated the format of the references in the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02694 2025-03-07 cs.CL cs.AI 62%

HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly

Howard Yen, Tianyu Gao, Minmin Hou, Ke Ding, Daniel Fleischer, Peter Izsak, Moshe Wasserblat, Danqi Chen

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

Comments ICLR 2025. Project page: https://princeton-nlp.github.io/HELMET/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00596 2025-03-04 cs.CL cs.AI cs.CR 62%

BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge

Terry Tong, Fei Wang, Zhe Zhao, Muhao Chen

专题命中 RAG评测 :RAG(abstract);分类 cs.CL、cs.AI

Comments Published to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏