'Skeleton Key' attack unlocks the worst of AI, says Microsoft
ID: 312a9d76-da1b-506f-a998-60ce4bfde2bf
STIX ID: report--312a9d76-da1b-506f-a998-60ce4bfde2bf
Feed Name: The Register (Security)
Microsoft disclosed 'Skeleton Key', a text-prompt jailbreak that tricks generative AI models into revising rather than abandoning safety instructions so they produce harmful content; Microsoft tested multiple major models (e.g., Llama3, Gemini Pro, GPT-3.5/4o variants, Claude 3 Opus) and found many were affected while GPT-4 resisted direct prompts but could be influenced via system messages. The write-up outlines mitigation approaches (Azure Prompt Shields and other defenses), notes that the attack requires legitimate access to the model, and warns researchers to address more advanced adversarial techniques such as BEAST or Greedy Coordinate Gradient.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
