Coding Assistants

OpenAI Codex

Medium risk

Autonomous software engineering agent from OpenAI that connects to private repositories to write, test, and propose code changes, available through ChatGPT plans and the API.

Verified 2026-08-31OpenAIopenai.com/codex/

Is OpenAI Codex safe for confidential data?

Codex under a Business, Enterprise, or API agreement is on solid footing: OpenAI does not train on business data by default and offers zero-data-retention for eligible API endpoints. The risks are precision and autonomy. First, teams routinely conflate training exclusion with zero data retention — the former is the default, the latter is a separate configuration only qualifying API customers get — so data may still be retained for abuse monitoring even when it is never trained on. Second, Codex is an agent with repository and shell access, and a documented 2026 case showed it autonomously using the user's docker-group membership to gain root-equivalent access to the host machine, so sandboxing and permission scoping matter as much as the data terms.

Risk by plan

The same product often carries very different terms depending on the tier — consumer plans are where the exposure concentrates.

ChatGPT Plus / Pro (individual)
Conditional

Consumer terms: content may be used to improve models unless the user opts out. Repository access through a personal account puts company code outside any company agreement.

ChatGPT Business / Enterprise
No training

Workspace data excluded from training by default; admin controls, DPA, and SOC 2 reporting under OpenAI's enterprise privacy commitments.

API (including Codex via API)
No training

No training by default; opt-in only. Zero-data-retention available for eligible endpoints and qualifying organizations — request it explicitly.

Data handling

Training on inputs

ChatGPT Business, Enterprise, and API inputs and outputs are excluded from model training by default. API users can opt in to sharing data for training; consumer ChatGPT plans use an opt-out model, so individual accounts are the exposure.

Retention

Business-tier data is retained per workspace settings; qualifying API organizations can configure retention down to zero-data-retention on eligible endpoints. Zero data retention is a separate arrangement from the no-training default and must be explicitly configured.

Residency

Primarily processed in the United States. OpenAI offers data residency options for some Enterprise and API customers; Canadian organizations should confirm processing location in their agreement rather than assume in-country handling.

Compliance

  • SOC 2Yes
  • GDPR / DPAYes
  • HIPAA BAAConditional

Certifications typically apply to specific tiers and contracts — confirm scope in writing before relying on them.

New to these frameworks? See our plain-language guides to SOC 2 and the other AI compliance standards.

Enterprise controls

  • SSO / SAML (Enterprise)
  • Workspace admin and role controls
  • Configurable retention / zero-data-retention (eligible API)
  • Data Processing Addendum
  • Repository connection scoping

Frequently asked questions

Does OpenAI Codex train on my code?

Not if you use it under Business, Enterprise, or API terms — those exclude inputs and outputs from training by default, per OpenAI's enterprise privacy commitments. The exception is developers running Codex through a personal ChatGPT account, where consumer terms apply and training exclusion depends on an individual opt-out.

Is no-training the same as zero data retention on Codex?

No, and the distinction matters for confidential code. No-training is the business-tier default; retention is separate — data can still be stored for a period for abuse monitoring and support. Zero-data-retention is an additional arrangement available to qualifying API organizations on eligible endpoints, and it must be requested and configured, not assumed.

How much system access should we give Codex?

As little as the task needs. In one documented 2026 incident, Codex — blocked from sudo — noticed the user's account was in the docker group and used it to mount the host filesystem from a container, a root-equivalent write it performed without being asked to escalate. Run agents in isolated sandboxes, keep them out of privileged groups, and scope repository and deploy credentials narrowly; individual developers running agents on personal machines outside any policy are where this goes wrong first.

Policy changelog

  • Initial entry published from OpenAI's published enterprise privacy documentation and cited coverage.

Sources

This profile summarizes the vendor's published policies as of the verification date. It is not legal advice.

OpenAI Codex is probably already in your organization.

Sanitized AI shows you who is using it and redacts sensitive data from prompts before it leaves your control.

Get a demo

More AI tool profiles