AI failed to properly patch software flaws 74% of the time, 1Password's study warns
ID: 177d59d0-5bfd-5076-a181-54735ad3a563
STIX ID: report--177d59d0-5bfd-5076-a181-54735ad3a563
Feed Name: ZDNet Security
1Password's Off-By-1-Labs evaluated frontier LLMs (including Claude and a Codex-based agent) on patch generation for six recently disclosed open-source CVEs, producing 6,080 patch attempts; only 26% were suitable, 21% fixed the bug but altered application behavior, and 53.9% failed or introduced new defects. The study coins these problematic outputs “FLAWED” (Fix-Like Artifacts With Embedded Defects), highlights common failure modes, and releases tooling (FLAWED) for researchers; it concludes LLMs currently help triage and vulnerability discovery but are not reliable automated patchers and require human oversight.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
