Microsoft’s Solution to Generative AI Jailbreak Attacks: All You Need to Know

3 min read
Co-founder of JustAINews
Share
Key Points
  • Recent criticism highlights how AI models can be exploited by malicious actors using techniques like jailbreaks.
  • Skeleton Key, a newly exposed jailbreak method, affects popular AI chatbots, including OpenAI's ChatGPT and Google's Gemini.
  • Microsoft has implemented Prompt Shields to protect Azure AI models, emphasizing the importance of multilayered security measures.
Credits: Ashkan Forouzani / Unsplash

One of the most emotionally charged criticisms of artificial intelligence is the concern that these technologies can be exploited by malicious actors for harmful purposes or simply for amusement.

A common method used to achieve this is through AI jailbreaks, a type of hacking aimed at bypassing the ethical safeguards of AI models. Recently, Microsoft exposed a new jailbreak method called Skeleton Key, which has been effective against several leading AI chatbots, including OpenAI's ChatGPT, Google's Gemini, and Anthropic's Claude.

Understanding the Skeleton Key

The Skeleton Key method involves using a multi-step tactic to make the AI model disregard its safety mechanisms. Once bypassed, the model cannot differentiate between harmful and genuine requests. This comprehensive bypass ability has earned it the name Skeleton Key.

“In bypassing safeguards, Skeleton Key allows the user to cause the model to produce ordinarily forbidden behaviors, which could range from production of harmful content to overriding its usual decision-making rules.” 

Mark Russinovich, Chief Technology Officer of Microsoft Azure

A representation of this technique shows how a user issues a Skeleton Key prompt that overrides system messages, tricking the AI into generating content it normally wouldn’t. The threat depends on the attacker already having valid access to the AI system. By avoiding safeguards, the attacker can manipulate the model to generate harmful material or override normal decision-making protocols.

Source: Microsoft

Attack Flow: How Skeleton Key Works

The Skeleton Key attack entails instructing the model to modify its behavior protocols.

Instead of denying requests for potentially harmful details, the model is encouraged to issue a warning yet still fulfill the request. 

This attack is classified as Explicit: forced instruction-following.

For instance, a model might refuse to generate instructions for making a Molotov cocktail. Yet, if the requester claims it’s for research purposes and asserts their expertise in safety and ethics, the model might be persuaded to comply. The outcome would be the model providing the requested content with a disclaimer, as depicted in the screenshot below.

Models Affected by Skeleton Key

During tests conducted from April to May 2024, the Skeleton Key attack was effective on a variety of generative AI models, including:

  • Meta Llama3-70b-instruct (base)
  • Google Gemini Pro (base)
  • OpenAI GPT 3.5 Turbo (hosted)
  • OpenAI GPT 4.0 (hosted)
  • Mistral Large (hosted)
  • Anthropic Claude 3 Opus (hosted)
  • Cohere Commander R Plus (hosted)

During tests, various tasks concerning risk and safety were evaluated, such as content involving explosives, bioweapons, political discussions, self-harm, racism, drugs, graphic violence, and more. All examined models complied with the jailbreak requests, though they prefixed their outputs with a user-requested warning.

Mitigation and Protection Guidance

Microsoft has introduced updates to the LLM technology powering its AI solutions, including Copilot AI assistants, to thwart Skeleton Key attacks. To protect AI models against this threat, organizations should consider the following strategies:

  • Input Filtering: Azure AI Content Safety can identify and block inputs with malicious intent aimed at initiating jailbreaks.
  • System Message Design: Proper engineering of system messages guides LLMs in appropriate conduct while safeguarding against attempts to override security.
  • Output Filtering: Post-processing filters can detect and block harmful content generated by the model.
  • Abuse Monitoring: Tools like Microsoft Azure OpenAI Service abuse monitoring can trace recurring content patterns and identify safety violations.

Microsoft integrates these protections into its Azure AI and Microsoft Security services, offering customers a robust framework for safeguarding AI-driven applications.

Subscribe to JustAINews!

Get the industry's biggest AI news straight to your inbox.
Subscribe

Related posts

© Copyright 2025 - Just AI News - All Rights Reserved
linkedin facebook pinterest youtube rss twitter instagram facebook-blank rss-blank linkedin-blank pinterest youtube twitter instagram