Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
细粒度微调在激活差异中留下明显可读的痕迹
机构 * EPFL(苏黎世联邦理工学院) ; Ecole Normale Supérieure Paris-Saclay(巴黎-萨克雷高等师范学校) ; Université Paris-Saclay(巴黎-萨克雷大学) ; Harvard University(哈佛大学) ; MATS
专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI
AI总结 研究发现狭窄微调会在激活中留下明显痕迹,揭示了微调领域偏见,并警告了使用此类模型进行广泛微调研究的局限性。
Comments ICLR 2026