From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
从行为到理解:LLM代理中时间概念的符合可解释性
机构 * Georgia State University(佐治亚州立大学) ; Computer Science Lab, SRI(SRI计算机科学实验室) ; Army Cyber Institute(陆军网络学院) ; United States Military Academy(美国军事学院) ; University of Florida(佛罗里达大学)
AI总结 本文提出通过符合视角解释LLM代理时间概念演变的框架,结合逐步奖励建模与符合预测,识别时间概念的潜在方向,实验表明这些概念可线性分离,为LLM代理提供可靠故障检测和干预方法。
Comments Accepted at the Mechanistic Interpretability Workshop, 43rd International Conference on Machine Learning, Seoul, South Korea, 2026