Geometric Deviation as an Unsupervised Pre-Generation Reliability Signal: Probing LLM Representations for Answerability
几何偏差作为一种无监督预生成可靠性信号:探测大语言模型表示的可回答性
机构 * University of Southern California(南加州大学)
AI总结 研究能否通过测量隐藏状态与可回答参考集的偏差,利用表示几何提供预生成信号。在三个模型和三种提示形式上实验,发现几何主要编码任务形式,数学提示中有可区分性,代码提示有部分泛化,该信号在早期层出现。
Comments Accepted to TrustNLP 2026 (ACL Workshop). 11 pages, 3 figures, 3 tables
Journal ref In Proceedings of the 6th Workshop on Trustworthy NLP (TrustNLP 2026), pp. 353-363, 2026