Case Study: Fine-tuning Small Language Models for Accurate and Private CWE Detection in Python Code
案例研究:微调小型语言模型以在Python代码中实现准确且隐私的CWE检测
机构 * Institute of Information and Communication Technology, Bangladesh University of Engineering Technology(孟加拉工程科技大学信息与通信技术研究所) ; Hajee Mohammad Danesh Science and Technology University(海杰莫哈默德丹什科学与技术大学)
专题命中 其他AI编程 :code model(abstract);分类 cs.AI
AI总结 本文研究了微调小型语言模型用于在Python代码中准确检测CWEs的可能性,通过半监督方法生成数据并人工审核,最终在测试集上实现了99%的准确率和99.04%的F1分数,展示了其在隐私保护下的高效安全分析潜力。
Comments 11 pages, 2 figures, 3 tables. Dataset available at https://huggingface.co/datasets/floxihunter/synthetic_python_cwe. Model available at https://huggingface.co/floxihunter/codegen-mono-CWEdetect. Keywords: Small Language Models (SLMs), Vulnerability Detection, CWE, Fine-tuning, Python Security, Privacy-Preserving Code Analysis