AI Security Glossary

The vocabulary of AI risk, in plain language

42 terms across AI fundamentals, threats, privacy, and governance — each defined the way you'd explain it to a colleague, and linked to the tool profiles and compliance guides where it matters.

Core AI Concepts

AI agent

An AI agent is a model given tools and autonomy — it browses, runs code, edits files, and takes multi-step actions, expanding both capability and blast radius.

Context window

The context window is the maximum amount of text an AI model can consider in one exchange — and the reason tools encourage pasting whole documents.

Embedding

An embedding is a numeric representation of text's meaning, used for semantic search and RAG — derived from your content, and sensitive in proportion to it.

Fine-tuning

Fine-tuning is additional training that adapts a base AI model to a specific task or organization using a curated dataset — a dataset that often contains the sensitive examples.

Inference

Inference is a trained AI model generating output from your input — the processing step every prompt goes through on the provider's servers.

Large language model (LLM)

A large language model is an AI system trained on massive text corpora to generate and understand language — the technology behind ChatGPT, Claude, Gemini, and most workplace AI tools.

Prompt

A prompt is everything submitted to an AI model in a request — the typed question plus any pasted text, uploaded files, and context the tool adds automatically.

Retrieval-augmented generation (RAG)

RAG is an architecture where an AI retrieves relevant documents from a knowledge base and includes them in the prompt — grounding answers in your data, and exposing your data to the prompt.

System prompt

A system prompt is the hidden instruction layer that configures an AI's behavior before the user types anything — the rules, persona, and constraints the deployer sets.

Token

A token is the unit of text an AI model processes — roughly a short word or word-fragment. Pricing, context limits, and retention all count in tokens.

Threats & Failure Modes

BYOAI (Bring Your Own AI)

BYOAI is employees bringing personal AI accounts and tools into work — the consumer-tier subscriptions and free chatbots that sit outside corporate contracts and controls.

Data leakage (AI)

AI data leakage is confidential information leaving an organization through prompts, uploads, or AI outputs — client names in a chatbot, source code in a coding assistant, a customer file in an AI answer engine.

Deepfake

A deepfake is AI-generated audio, video, or imagery that convincingly imitates a real person — used in fraud, impersonation, and disinformation.

Hallucination

A hallucination is confident but false output from an AI model — invented facts, citations, numbers, or case law presented as real.

Jailbreak (AI)

A jailbreak is a prompt crafted to make an AI model ignore its safety rules and produce output it was trained to refuse.

Prompt injection

Prompt injection is an attack that hides instructions in content an AI processes — a document, web page, or email — so the model follows the attacker's commands instead of the user's.

Red teaming (AI)

AI red teaming is adversarial testing of AI systems — probing for jailbreaks, prompt injection, data leaks, and harmful outputs before attackers or users find them.

Shadow AI

Shadow AI is employees using AI tools their organization hasn't approved or can't see — personal ChatGPT accounts, browser extensions, meeting bots — creating data-leakage and compliance risk outside IT's controls.

Social engineering (AI-enabled)

AI-enabled social engineering uses generative tools to scale deception — fluent phishing in any language, deepfaked voices, and chatbots that hold convincing fraudulent conversations.

Voice cloning

Voice cloning is AI replication of a specific person's voice from sample audio — a biometric-data processing activity with consent obligations, and a fraud vector when abused.

Privacy & Data Protection

Anonymization

Anonymization irreversibly removes the link between data and identifiable people. Done properly, the data exits privacy-law scope — a high bar that casual name-stripping doesn't meet.

Biometric data

Biometric data measures unique physical traits — faces, voices, fingerprints. AI tools that clone voices or generate avatars process it, triggering the strictest consent rules.

Data residency

Data residency is where data is physically stored and processed. AI tools typically process prompts in the vendor's home cloud — often the US — unless an enterprise tier says otherwise.

Data sovereignty

Data sovereignty is the principle that data is subject to the laws of the country where it's located or controlled — the legal layer above data residency.

De-identification

De-identification is the umbrella practice of stripping identifying elements from data — spanning redaction, pseudonymization, and anonymization, each with different legal weight.

PHI (Protected Health Information)

PHI is health information tied to an identifiable person — diagnoses, treatments, health numbers. It carries the strictest handling rules of any routine business data.

PII (Personally Identifiable Information)

PII is any information that identifies a person — names, emails, ID numbers, and combinations that single someone out. Most privacy law obligations attach to it.

Pseudonymization

Pseudonymization replaces identifiers with tokens or aliases that can be reversed with a separately-held key — reducing exposure while the data remains personal information.

Redaction (automated)

Automated redaction detects and removes sensitive data — names, numbers, identifiers — from text before it leaves, replacing it with placeholders so the task still works.

Zero data retention (ZDR)

Zero data retention is a provider commitment not to store prompts or outputs after the request completes — the strongest data-handling posture an AI vendor offers.

Governance & Compliance

AI acceptable-use policy (AUP)

An AI acceptable-use policy sets the rules for AI at work: approved tools, prohibited data, verification duties, and consequences — the written half of AI governance.

AI DLP

AI DLP is data-loss prevention for AI interactions — detecting and redacting sensitive data in prompts and uploads before they reach chatbots and assistants, where classic DLP is blind.

AI governance

AI governance is the policies, controls, and oversight an organization applies to AI use — who may use what, with which data, under whose review, with what evidence.

AI risk assessment

An AI risk assessment systematically evaluates an AI tool or use case before adoption — data flows, training terms, compliance scope, misuse potential — and documents the verdict.

Audit trail (AI usage)

An AI audit trail is the logged record of AI interactions — who used which tool, what was caught, what was allowed — the evidence layer regulators and clients ask for.

BAA (Business Associate Agreement)

A BAA is the HIPAA contract required before any vendor touches protected health information — including AI vendors. Without one, PHI in a prompt is a violation.

DPA (Data Processing Agreement)

A DPA is the contract governing how a vendor processes personal data on your behalf — purposes, security, subprocessors, breach duties. No DPA, no regulated data.

Human review (of AI conversations)

Human review is vendor staff or contractors reading sampled user conversations for quality and safety — meaning a person, not just a model, may see what was pasted.

Sanctioned vs. unsanctioned AI

Sanctioned AI is the tooling an organization approved and contracts for; unsanctioned AI is everything else employees actually use. Governance lives in the gap.

Security awareness training (AI)

AI-focused security awareness training teaches employees what not to paste and why — and works best paired with a control that fires at the moment of the paste.

Subprocessor

A subprocessor is a vendor's vendor — the model providers and clouds your AI tool sends data to. Your data's real journey is the subprocessor list.

Training on inputs

Training on inputs is a vendor using your prompts and uploads to improve its models — the single most consequential line in any AI tool's terms, and it varies by tier.

From vocabulary to visibility.

Knowing what shadow AI is doesn't show you yours. Sanitized AI inventories the AI tools in use and redacts sensitive data before it leaves.

Get a demo