arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

2026-07-07 至 2026-07-07 共收录 9 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 检索器与排序 5 篇

2508.02296 2026-07-07 cs.CL cs.IR 版本更新 90%

Knowing When Not to Answer: Lightweight KB-Aligned OOD Detection for Safe RAG

知道何时不回答:用于安全检索增强生成的轻量级知识库对齐的域外检测

Ilias Triantafyllopoulos, Renyi Qu, Salvatore Giorgi, Brenda Curtis, Lyle H. Ungar, João Sedoc

机构 * New York University(纽约大学) National Institute on Drug Abuse(国家成瘾医学研究所) University of Pennsylvania(宾夕法尼亚大学) Microsoft(微软公司)

专题命中 检索器与排序 :RAG(title,summary_cn);retrieval-augmented generation(abstract);dense retrieval(abstract);分类 cs.IR、cs.CL

AI总结 研究用于检索增强生成(RAG)系统的轻量级、知识库对齐的域外检测,通过PCA处理知识库嵌入,用方差保留或t检验排序选择子空间评分查询,评估规则和分类器,发现低维检测器性能好且更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29580 2026-07-07 cs.CL 版本更新 79%

MAM-AI: An On-Device Medical Retrieval-Augmented Generation System for Nurses and Midwives in Zanzibar

MAM-AI:面向桑给巴尔护士和助产士的设备端医疗检索增强生成系统

Yi Ren

机构 * École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院洛桑分校)

专题命中 检索器与排序 :retrieval-augmented generation(title);retriever(abstract);分类 cs.CL

AI总结 针对撒哈拉以南非洲地区护士助产士难以获取权威指南的问题,提出完全运行于安卓设备上的医疗问答系统MAM-AI,采用300M嵌入模型检索87份指南文档,并用4B int4生成器离线生成带引用的回答,评估发现小生成器在安全性和有用性间存在权衡,通过优化提示词降低回避率。

Comments 38 pages. Video demo: https://www.youtube.com/watch?v=M_Kruluel28 ; browser demo, code, models, and benchmarks linked in the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29467 2026-07-07 cs.CL cs.IR 版本更新 76%

mamabench and mamaretrieval: Benchmarks for Evaluating Medical Retrieval-Augmented Generation in Maternal, Neonatal, and Reproductive Health

mamabench 和 mamaretrieval:评估孕产妇、新生儿和生殖健康领域医学检索增强生成的基准

Yi Ren

机构 * École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院洛桑校区)

专题命中 检索器与排序 :retrieval-augmented generation(title);分类 cs.IR、cs.CL

AI总结 针对助产士咨询的孕产妇、新生儿和生殖健康问题,构建了包含25,949个问题的QA基准mamabench和基于3,185个查询的段落级相关性基准mamaretrieval,采用分级相关标注和标签质量审计。

Comments 13 pages, 3 tables. Datasets and construction code linked in the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02950 2026-07-07 cs.CL cs.AI cs.HC 版本更新 62%

Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues

固定合成日语咨询对话中的结构化提示与自动评估

Keita Kiuchi, Yoshikazu Fujimoto, Hideyuki Goto, Tomonori Hosokawa, Makoto Nishimura, Yosuke Sato, Izumi Sezai, Tomohiro Inoue

机构 * Japan National Institute of Occupational Safety and Health(日本国立职业安全卫生研究所) Kaze To Taiyo(凯泽・太阳) Saga Occupational Health Association(Saga职业健康协会) Department of Pharmacy, Zikei Hospital/Zikei Institute of Psychiatry(药剂科,Zikei医院/Zikei精神医学研究所) Department of Medical Welfare, Suzuka University of Medical Science(医疗福祉科, Suzuka医科大学) Graduate School of Human Sciences, Ritsumeikan University(人类科学研究生院,立命馆大学) Faculty of Nursing, National Defense Medical College(护理学部,国家防卫医疗大学) Support Center for Students with Disabilities, Aoyama Gakuin University(残疾学生支持中心,上智大学)

专题命中 检索器与排序 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 研究通过人工智能生成的18份固定日语咨询对话记录,对比GPT - minimal、GPT - SMDP和Claude - SMDP三种情况,咨询专家与新的语言模型分别评级,发现SMDP对话在多方面获更高专家评级,语言模型评级可重复但偏宽松。

Comments 59 pages, 2 figures, 30 tables; supplemental material included; data and code at https://doi.org/10.5281/zenodo.21182321; preregistration at https://doi.org/10.17605/OSF.IO/VU286

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07285 2026-07-07 cs.CL cs.CY 版本更新 57%

Why teaching resists automation in an AI-inundated era: Human judgment, non-modular work, and the limits of delegation

为何在人工智能泛滥的时代教学抵抗自动化:人类判断、非模块化工作与委托的局限

Songhee Han

机构 * Florida State University(佛罗里达州立大学)

专题命中 检索器与排序 :retrieval-augmented generation(abstract);分类 cs.CL

AI总结 本文探讨了在人工智能普及背景下,教学工作难以自动化的原因,指出教学本质上具有解释性、关联性和专业判断,无法被完全自动化或委托给技术。

Comments Revised version; accepted for publication in TechTrends

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 向量检索 1 篇

2510.00566 2026-07-07 cs.LG cs.AI cs.DB 版本更新 73%

Panorama: Fast-Track Nearest Neighbors

全景图:快速跟踪最近邻

Vansh Ramani, Alexis Schlomer, Akash Nayar, Sayan Ranu, Jignesh M. Patel, Panagiotis Karras

机构 * Carnegie Mellon University(卡内基梅隆大学) IIT Delhi(德里印度理工学院) UCPH(乌兹堡大学) Indian Institute of Technology Delhi(德里印度理工学院) University of Copenhagen(哥本哈根大学)

专题命中 向量检索 :retrieval-augmented generation(abstract);RAG(abstract);分类 cs.AI、cs.DB

AI总结 针对高维神经嵌入的近似最近邻搜索,提出PANORAMA技术,利用嵌入固有频谱衰减加速验证,通过PCA压缩信号能量,增量评估候选距离,解决与乘积量化假设冲突问题,在FAISS库中优化,实现加速。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 知识库问答 1 篇

2605.05409 2026-07-07 cs.AI cs.CL 版本更新 90%

Agentic Retrieval-Augmented Generation for Financial Document Question Answering

代理检索增强生成用于财务文档问答

Yang Shu, Yingmin Liu, Zequn Xie

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

专题命中 知识库问答 :RAG(summary_cn,abstract);retrieval-augmented generation(title,abstract);retriever(abstract);分类 cs.CL、cs.AI

AI总结 本文提出FinAgent-RAG框架,通过迭代检索-推理循环和自我验证,提升金融文档问答的精度。引入对比金融检索器、程序化思维模块和自适应策略路由,实验表明在三个基准数据集上均取得显著效果,准确率提升5.62-9.32个百分点。

Comments This paper is withdrawn due to significant methodological errors in the experimental design that fundamentally affect the validity of the results. The errors are not correctable within the current framework, and the conclusions can no longer be supported. We apologize for any inconvenience caused to readers

详情

展开后加载摘要…

URL PDF HTML 收藏

4. RAG评测 2 篇

2408.13378 2026-07-07 cs.AI cs.CL cs.IR cs.LG q-bio.QM 版本更新 80%

DrugAgent: Reliable Multi-Agent Integration of Conflicting Biomedical Evidence for Drug-Target Interaction Assessment

DrugAgent:用于药物-靶点相互作用评估的冲突生物医学证据的可靠多智能体整合

Yoshitaka Inoue, Tianci Song, Xinling Wang, Rui Kuang, Tianfan Fu, Augustin Luna

机构 * Department of Computer Science and Engineering, University of Minnesota(计算机科学与工程系,明尼苏达大学) Computational Biology Branch, National Library of Medicine(国家医学图书馆计算生物学分支) Khoury College of Computer Sciences, Northeastern University(东北大学计算机科学学院) State Key Laboratory for Novel Software Technology at Nanjing University, School of Computer Science, Nanjing University(南京大学新型软件技术国家重点实验室,南京大学计算机科学学院) Developmental Therapeutics Branch, National Cancer Institute(国家癌症研究所发育治疗分支)

专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL、cs.AI

AI总结 研究药物-靶点相互作用评估中整合异构数据问题,核心方法是基于大语言模型的多智能体系统DrugAgent,贡献是能进行异构证据支持的DTI评估,补充独立DTI预测,还提供相关策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06142 2026-07-07 cs.CL cs.AI 版本更新 79%

IRC-Bench: Recognizing Entities from Contextual Cues in First-Person Reminiscences

IRC-Bench: 从第一人称回忆中的上下文线索识别实体

Yehudit Aperstein, Eden Moran, Alexander Apartsin

机构 * Intelligent Systems, Afeka Academic College of Engineering(阿法卡学术工程学院智能系统) School of Computer Science, Faculty of Sciences, Holon Institute of Technology(霍隆理工学院计算机科学学院)

专题命中 RAG评测 :RAG(abstract,abstract_cn);dense retrieval(abstract);分类 cs.CL、cs.AI

AI总结 本文提出IRC-Bench基准,用于评估回忆文本中隐式实体识别任务,通过对比本地提及与分散叙述证据,探讨非局部性挑战,测试多种模型配置。

Comments 36 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏