Prefill Awareness in Large Language Models
大型语言模型中的预填充感知
机构 * Constellation University of Wisconsin-Madison(威斯康星大学麦迪逊分校星座研究所) ; Constellation Georgia Institute of Technology(佐治亚理工学院星座研究所) ; UK AI Security Institute(英国人工智能安全研究所)
AI总结 研究大型语言模型能否识别并响应其助手消息被预填充或篡改,发现前沿模型具有显著预填充感知能力,可能影响安全评估方法。
Comments Submitted to NeurIPS 2026