Codex Security
OpenAI's application security agent for the terminal and CI
About Codex Security
Codex Security is OpenAI's application security agent, shipped as the Apache-2.0 npm package @openai/codex-security containing both a CLI and a TypeScript SDK. It scans a repository for likely vulnerabilities, validates candidate findings before it reports them, and can produce bounded patches for findings you approve. The same scanner is available four ways: as a plugin inside the ChatGPT desktop app, from your terminal, through the SDK, and as Codex Security cloud against connected GitHub repositories.
The validation pass is the part that distinguishes it from running a model over your source. Security scanning fails on noise long before it fails on coverage, and a tool that reports fifty plausible issues of which three are real is worse than no tool, because the triage cost lands on the people least able to absorb it. Codex Security builds a repository-specific threat model, checks likely vulnerabilities against real code context rather than generic signatures, and validates high-signal issues in an isolated environment before surfacing them, with the evidence attached. Two caveats belong up front. The package is at an early version and moving fast, and there is an access gate: the README states that some cybersecurity requests and protected findings require approval through Trusted Access for Cyber, and the documentation adds that running scans requires Codex Security access, with a Trusted-Access-verified account recommended. The exact boundary of that gate is not published, so budget for the possibility that a scan you want is out of scope. Codex Security cloud, the GitHub-connected half, is described by OpenAI as a research preview.
Install the package with npm (Node 22.13.0 or later and Python 3.10 or later are required), run codex-security login, then codex-security scan /path/to/directory. In CI you skip the login and set OPENAI_API_KEY instead. A scan discovers repositories, tracks findings across runs so results are comparable over time, and records false-positive feedback so a dismissed finding stays dismissed.
From there it fans out. Bulk campaigns run from a CSV inventory and resume where they stopped. CI runs review pull-request changes, preserve artifacts, upload SARIF and enforce a severity policy. The TypeScript SDK exposes the same scanner programmatically, with modes, worker counts, subagent counts and time limits as parameters, so you can build scanning and progress reporting into your own developer tooling.
- •Validation Before Reporting - Candidate findings are checked in an isolated environment against a repository-specific threat model before they reach you, with validation evidence attached, which is aimed squarely at the false-positive problem
- •Four Surfaces, One Scanner - A ChatGPT desktop plugin with a Security workbench for scans, findings and repositories; a terminal CLI; a TypeScript SDK; and Codex Security cloud for connected GitHub repos
- •Provider Independence - Runs against OpenAI, Amazon Bedrock, OpenRouter or Fireworks by setting the relevant key and selecting a model, so it is not tied to one inference vendor
- •CI Integration with SARIF - Review changes on a pull request, preserve artifacts, upload SARIF and apply a severity policy as a merge gate
- •Cost Controls -
--max-costsets an estimated ceiling in USD; when a deep scan reaches it, the CLI saves a report with partial coverage and exits with code 2 rather than silently continuing - •Bulk and Containerized Scans - A Docker Compose configuration runs campaigns across many repositories, with a separate findings service (in preview) that stores findings and embeddings in SQLite and groups duplicates by embedding similarity
- •Fix and Verify - Bounded patches for approved findings, with verification, rather than a report that stops at the diagnosis
Security engineers who own an application security program and are drowning in scanner output, and engineering teams who want a security gate on pull requests without buying a platform. The SARIF output and severity policy make it a drop-in for an existing CI pipeline, and the SDK is there for anyone building security review into their own internal tooling. Teams already using Codex in the ChatGPT desktop app get the workbench with no extra setup.
It is a weaker fit in two situations. If you need a guaranteed, contractually supported scanner for a compliance audit, this is a pre-1.0 package with a cloud component in research preview and an access program that may gate what you can run. And if your work involves offensive security or sensitive cyber requests, expect to go through Trusted Access for Cyber first rather than assuming an install is enough.














