Sequential Correlations Change In-Context Learning: Effective Context Length and Architectural Mismatch
序列相关性改变上下文学习:有效上下文长度与架构失配
机构 * John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保罗森工程与应用科学学院) ; Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究所) ; Society of Fellows and Center for Brain Science, Harvard University(哈佛大学 fellows 会与脑科学中心)
AI总结 针对现有ICL理论多聚焦独立样本提示的局限,研究序列关联数据下的ICL特性,基于线性注意力构建可解模型,发现关联提示会改变ICL有效样本量及最优适配注意力架构。