logo

Show top LLMs some code and they'll merrily add in the bugs they saw in training

ID: ba5e5ca7-348f-52c0-a382-56aa73ca7aa5

STIX ID: report--ba5e5ca7-348f-52c0-a382-56aa73ca7aa5

Feed Name: The Register (Security)

Date Published: 2025-03-19

Date Updated: 2026-04-26

Author: Thomas Claburn

...
...

Researchers evaluated multiple LLMs (GPT-4o, GPT-3.5, GPT-4, CodeLlama-13B, Gemma-7B, StarCoder2-15B, CodeGEN-350M, and DeepSeek R1) on completing bug-prone code from the Defects4J dataset and found the models frequently replicate existing bugs rather than fix them, resulting in significantly lower accuracy than normal code completion (e.g., GPT-4 at 12.27% vs 29.85%), with a large portion of buggy outputs identical to historical bugs (up to 82.61% for GPT-4o); the authors recommend improving semantic understanding, error detection/handling, post-processing, and IDE integration.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.