Pretraining Exposure Explains Popularity Judgments in Large Language Models
预训练暴露解释了大型语言模型中的流行度判断
机构 * University of Innsbruck(因斯布鲁克大学)
专题命中 预训练与数据 :large language model(title,abstract);language model(title,abstract);pretraining(title,abstract);LLM(abstract,abstract_cn)
AI总结 研究通过分析预训练数据中的暴露统计,发现大型语言模型对流行实体的偏好主要受预训练暴露影响,而非外部流行度信号,揭示了数据暴露在驱动流行度偏差中的核心作用。
Comments Accepted at SIGIR 2026
Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026)