logo

NDSS 2025 – A Comparative Evaluation Of Large Language Models In Vulnerability Detection

ID: 29850c48-d04a-569c-8e42-98174a4bd971

STIX ID: report--29850c48-d04a-569c-8e42-98174a4bd971

Feed Name: Security Boulevard

Date Published: 2026-03-03

Date Updated: 2026-04-22

Author: Marc Handelman

...
...

This NDSS conference paper evaluates multiple large language models (LLaMA-2, CodeLLaMA, LLaMA-3, Mistral, Mixtral, Gemma, CodeGemma, Phi-2, Phi-3, GPT-4) for automated vulnerability detection, primarily in Java with additional C/C++ tests. The study measures runtime and detection accuracy in zero-shot and few-shot settings, examines effects of model parameters (including quantization and context window size), and finds that while some models (e.g., Gemma, LLaMA-2) can perform well, performance varies widely by model, language, and configuration, with limitations remaining for generalized, reliable vulnerability identification.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.