Lean and Mean: How We Fine-Tuned a Small Language Model for Secret Detection in Code
ID: 6d929481-2eeb-52ba-bb7b-69e9daf60d8d
STIX ID: report--6d929481-2eeb-52ba-bb7b-69e9daf60d8d
Feed Name: Wiz Blog
Wiz presents a research effort to fine-tune a small language model (Llama 3.2 1B) for secret detection in code, addressing the scalability, cost, and privacy limitations of large LLMs and the high false positives of regex. Using a multi-agent LLM labeling pipeline, data filtration, LoRA adapters, and selective quantization, the model achieves ~82% recall and ~86% precision while running efficiently on CPUs. The work includes a validation framework, inference optimizations, and a phased production rollout that augments existing regex-based detection, with future plans to expand beyond code and enhance contextual risk assessment.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
