arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

University of Pennsylvania(宾夕法尼亚大学)

2026-03-30 至 2026-03-30 共收录 5
2601.13227 2026-03-30 cs.IR cs.AI

Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?

内部知识:RAG系统能从评估秘密中获得多少收益?

Laura Dietz, Bryan Li, Eugene Yang, Dawn Lawrie, William Walden, James Mayfield

机构 * University of New Hampshire(新罕布什尔大学) University of Pennsylvania(宾夕法尼亚大学) Human Language Technology Center of Excellence, Johns Hopkins University(约翰霍普金斯大学人类语言技术卓越中心)

AI总结 本文通过对比实验探讨RAG系统在评估秘密泄露时的评估风险,指出盲评和方法多样性的重要性。

Comments To appear in ECIR 2026, Lecture Notes in Computer Science, Volume 16483

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13222 2026-03-30 cs.IR cs.AI

Incorporating Q&A Nuggets into Retrieval-Augmented Generation

将问答 nugget 融入检索增强生成

Laura Dietz, Bryan Li, Gabrielle Liu, Jia-Huei Ju, Eugene Yang, Dawn Lawrie, William Walden, James Mayfield

机构 * University of New Hampshire(新罕布什尔大学) University of Pennsylvania(宾夕法尼亚大学) Yale University(耶鲁大学) University of Amsterdam(阿姆斯特丹大学) Human Language Technology Center of Excellence, Johns Hopkins University(约翰霍普金斯大学人类语言技术卓越中心)

AI总结 本文提出Crucible系统,通过构建问答nugget库提升生成质量,保留引用溯源,优于现有系统。

Comments To appear in the Proceedings of ECIR 2026, Lecture Notes in Computer Science, Volume 16484

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20093 2026-03-30 cs.CL cs.AI

Evaluating Neural Language Models as Cognitive Models of Language Acquisition

评估神经语言模型作为语言习得的认知模型

Héctor Javier Vázquez Martínez, Annika Lea Heuser, Charles Yang, Jordan Kodner

机构 * University of Pennsylvania(宾夕法尼亚大学) Stony Brook University(石溪大学)

AI总结 本文指出现有神经语言模型语法能力评估基准不够严谨,建议使用经过严格评估的语料库以更准确研究语言习得机制。

Comments To appear in the GenBench 2023 workshop proceedings, the first workshop on (benchmarking) generalisation in NLP. GenBench 2023 will be held at EMNLP 2023 on December 6, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25945 2026-03-30 eess.IV cs.CV

Adapting Segment Anything Model 3 for Concept-Driven Lesion Segmentation in Medical Images: An Experimental Study

为医学图像中的概念驱动病变分割适应Segment Anything Model 3:一项实验研究

Guoping Xu, Jayaram K. Udupa, Yubing Tong, Xin Long, Ying Zhang, Jie Deng, Weiguo Lu, You Zhang

机构 * The Medical Artificial Intelligence and Automation (MAIA) Laboratory, Department of Radiation Oncology, University of Texas Southwestern Medical Center(德克萨斯大学西南医学中心放射肿瘤学系医学人工智能与自动化(MAIA)实验室) Medical Image Processing Group, Department of Radiology, University of Pennsylvania(宾夕法尼亚大学放射学系医学图像处理组) Division of Digestive and Liver Diseases, Department of Internal Medicine, University of Texas Southwestern Medical Center(德克萨斯大学西南医学中心内科学系消化与肝脏疾病科)

AI总结 本文评估SAM3在多模态医学图像中的病变分割性能,通过几何框和概念提示提升鲁棒性,并展示其跨模态泛化能力。

Comments 31 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22498 2026-03-30 cs.RO

HELIOS: Hierarchical Exploration for Language-Grounded Interaction in Open Scenes

HELIOS:面向开放场景的语言引导交互的分层探索

Katrina Ashton, Chahyon Ku, Shrey Shah, Saumit Vedula, Tingrui Zhang, Wen Jiang, Kostas Daniilidis, Bernadette Bucher

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Michigan(密歇根大学)

AI总结 HELIOS通过分层场景表示和搜索目标,解决开放环境中语言指令与部分观测场景的语义关联及动态更新问题,实现高仿真实验和现实场景中的语言引导抓取放置任务。

详情

展开后加载摘要…

URL PDF HTML 收藏