Show top LLMs some code and they'll merrily add in the bugs they saw in training
ID: ba5e5ca7-348f-52c0-a382-56aa73ca7aa5
STIX ID: report--ba5e5ca7-348f-52c0-a382-56aa73ca7aa5
Feed Name: The Register (Security)
Researchers evaluated multiple LLMs (GPT-4o, GPT-3.5, GPT-4, CodeLlama-13B, Gemma-7B, StarCoder2-15B, CodeGEN-350M, and DeepSeek R1) on completing bug-prone code from the Defects4J dataset and found the models frequently replicate existing bugs rather than fix them, resulting in significantly lower accuracy than normal code completion (e.g., GPT-4 at 12.27% vs 29.85%), with a large portion of buggy outputs identical to historical bugs (up to 82.61% for GPT-4o); the authors recommend improving semantic understanding, error detection/handling, post-processing, and IDE integration.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
