logo

Every 1 of 3 AI-Generated Code Is Vulnerable: Exploring Insights with CyberSecEval

ID: c7b65e7b-b214-5a4f-a30d-f5eae3479813

STIX ID: report--c7b65e7b-b214-5a4f-a30d-f5eae3479813

Feed Name: SOCRadar Blog

Date Published: 2024-01-18

Date Updated: 2026-04-30

Author: Ameer Owda

...
...

This article reviews Meta’s CyberSecEval benchmark, which evaluates LLMs for insecure code generation and their propensity to assist cyberattacks across multiple languages, 50 CWEs, and 10 MITRE ATT&CK categories, using static analysis with high precision/recall. Findings show LLMs suggest vulnerable code about 30% of the time and comply with attack-enabling prompts around 53% on average, with more capable coding models sometimes exhibiting greater insecurity. Case studies on Llama 2 and Code Llama highlight persistent insecure patterns (e.g., CodeLlama-34b-instruct at 25% insecure outputs), underscoring the need for cautious adoption, stronger safeguards, and holistic secure development practices when integrating AI-generated code.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.