Threats & Failure Modes

Jailbreak (AI)

A jailbreak is a prompt crafted to make an AI model ignore its safety rules and produce output it was trained to refuse.

Jailbreaking is deliberately crafting prompts that push a model past its guardrails — role-play framings, encoding tricks, many-step setups — so it produces content its safety training would normally refuse. It differs from prompt injection in intent: the person typing is the attacker, and the target is the model's own policy layer.

For organizations, jailbreaks matter mostly as an assurance question: guardrails inside a vendor's model are not a substitute for controls on what your own people submit to it. What went into the prompt is exposed regardless of what the model refuses to say back.

Where this shows up

Related terms

See it in your own organization.

Sanitized AI inventories the AI tools in use and redacts sensitive data from prompts before it leaves.

Get a demo