A Validation-Gated Mechanistic Account of Suicidality Detection in LLMs
一种验证门控机制的自杀检测在LLMs中的解释
机构 * Intelligent Neuromorphic and Quantum Understanding for Innovative Research and Engineering (INQUIRE) Lab(智能神经形态与量子理解创新研究与工程实验室) ; School of Electrical and Computer Engineering(电气与计算机工程学院) ; University of Oklahoma(俄克拉荷马大学) ; School of Psychology(心理学学院) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 指令微调 :large language model(abstract);language model(abstract);分类 cs.CL
AI总结 提出验证门控框架,通过因果分析发现Llama-3.1-8B-Instruct在自杀检测中依赖语义特征而非关键词,且该特征在多个模型和数据集中一致出现。