GeoLAN: Geometric Learning of Latent Explanatory Directions in Large Language Models
GeoLAN:大型语言模型中潜在解释方向的几何学习
机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) ; Florida Institute for National Security Applied Artificial Intelligence Group(佛罗里达国家安全应用人工智能小组) ; University of Florida(佛罗里达大学)
专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);分类 cs.LG
AI总结 GeoLAN通过将token表示视为几何轨迹,结合Kakeya猜想相关条件,提升大语言模型的几何指标和公平性,尤其在中等规模模型中表现显著。