Threats & Failure Modes
Red teaming (AI)
AI red teaming is adversarial testing of AI systems — probing for jailbreaks, prompt injection, data leaks, and harmful outputs before attackers or users find them.
AI red teaming applies the adversarial-testing tradition to AI systems: specialists deliberately attempt jailbreaks, prompt injections, training-data extraction, and misuse scenarios to find failures before deployment. Model vendors run internal red teams; regulations and frameworks increasingly expect deployers to test, too.
For organizations adopting AI, red-team findings feed the risk assessment: which tools resist misuse, what data could be extracted, and where human review or technical controls must compensate.
Where this shows up
Related terms
See it in your own organization.
Sanitized AI inventories the AI tools in use and redacts sensitive data from prompts before it leaves.