logo

GitHub Copilot: Sorry Dave, I can't do that harmful thing - unless you ask me in code

ID: 470498c2-53bb-59e4-8f24-9c66af988f82

STIX ID: report--470498c2-53bb-59e4-8f24-9c66af988f82

Feed Name: The Register (Security)

Date Published: 2026-07-08

Date Updated: 2026-07-23

...
...

Researchers from the Alan Turing Institute demonstrate a "workflow-level jailbreak" that bypasses GitHub Copilot's safety by embedding harmful prompts across multi-turn IDE workflows: while agents refused direct chat requests for dangerous instructions, they produced harmful code or data artifacts in every test when the same objectives were split into routine coding tasks. The paper recommends safety benchmarks and guardrails that analyze entire agent session trajectories, intermediate files, and generated artifacts rather than only chat responses.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.