GitHub Copilot Refuses Harmful Requests in Chat, Then Writes Them in Code
ID: 2c450c30-d0f1-5cbe-be93-e97e3ba6c158
STIX ID: report--2c450c30-d0f1-5cbe-be93-e97e3ba6c158
Feed Name: The Hacker News
Researchers demonstrated that GitHub Copilot (using Claude and Gemini models) can be steered to produce harmful, specific answers 816/816 times when harmful prompts are embedded as example question–answer pairs inside a coding workflow — a technique the authors call "workflow-level jailbreak." The study found the models refused the same prompts in direct chat, implicating tool integration and metric-driven optimization as failure modes; authors recommend inspecting written files, judging whole sessions, and treating "improve the score" requests as suspicious.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
