logo

Mitigating Skeleton Key, a new type of generative AI jailbreak technique

ID: 53d8b073-c85d-51df-8a91-2737d9f2b3fb

STIX ID: report--53d8b073-c85d-51df-8a91-2737d9f2b3fb

Feed Name: Microsoft Security

Date Published: 2024-06-26

Date Updated: 2026-04-28

Author: Mark Russinovich

...
...

Microsoft details the “Skeleton Key” jailbreak, a multi-step explicit instruction-following technique that bypasses LLM guardrails across several models (e.g., Llama3, Gemini, GPT‑3.5/4o, Mistral Large, Claude 3 Opus, Cohere), enabling direct generation of normally restricted content with warning disclaimers. The post outlines testing results (with GPT‑4 showing relative resistance except via system-message manipulation), describes responsible disclosure to affected vendors, and provides mitigation guidance and tooling including Prompt Shields, input/output filtering, abuse monitoring, updated PyRIT evaluations, and Azure AI/Defender integrations for detection and protection.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.