AI EngineeringZero to ProductionHome·About·What’s new·Contact
Codex & OpenAI · Chapter O3

Codex & Agentic Development with OpenAI

Codex is OpenAI's coding agent that runs in your terminal — it reads your repo, plans, edits files, and runs commands in a loop. This is the OpenAI counterpart to the Claude Code chapter: how to install it, the agentic loop, and — most importantly — the sandbox and approval model that keeps it safe.

⏱️ ~1.5 hours🖥️ CLI tool🎯 Beginner→Tech-lead

Learning objectives

  • Explain what an agentic coding CLI is and when to reach for one.
  • Install Codex and run your first task, interactively and headless.
  • Reason about the agentic loop: plan → edit → run → iterate.
  • Configure sandbox modes and approval policy — the safety model.
  • Wire MCP servers and run Codex non-interactively in CI.
The commands on this page drive a real coding agent that reads, writes, and executes against your filesystem and shell. Run them only in a repo you can afford to have changed, start in the most restrictive sandbox, and review what the agent proposes before you widen its permissions.

1 · What an agentic coding CLI is essential

There's a leap between "the model writes code in a chat window" and "the model fixes the bug in your repo." In a chat window, you are the hands: you copy the suggestion, paste it into a file, run it, paste the error back. An agentic coding CLI closes that loop — it is the hands. It reads your files, proposes an edit, runs the test, sees the failure, and tries again, all without you shuttling text back and forth. Codex is OpenAI's version of that tool; it's the direct peer to Claude Code.

The mental model is a junior pair-programmer at your keyboard. You describe the goal ("make the failing test pass", "add a flag to this CLI"); the agent explores the codebase, forms a plan, makes changes, and runs commands to check its work. Crucially, it operates inside your environment — your files, your shell, your git — which is exactly what makes it powerful and exactly what makes the safety model (section 4) the most important part of this chapter.

The common mistake is treating it like autocomplete and letting it run unsupervised with full permissions on day one. It is far more capable than autocomplete and far more consequential: an agent that can run shell commands can also delete files or push commits. The professional posture is the opposite — start locked down, watch what it does, and widen trust only as you learn its behavior.

the agent is the hands — it reads, edits, runs, and observes in a loop read repo & plan edit files run commands tests · build observe iterate until the goal is met — bounded by your approvals & sandbox Read · edit · run · observe · repeat. The agent closes the loop you used to close by hand — but every edit and command is bounded by the sandbox and approval settings you'll configure in section 4.
🗺️ How to read this diagram
  • The first box is where the agent reads your repo and forms a plan from your goal.
  • The middle boxes are the actions — editing files and running commands (tests, builds) — the parts that touch your real environment.
  • The green box is observation: it reads the result and the dashed loop takes it back to planning until the goal is met.
  • The whole loop runs inside the limits you set — which is what section 4 is about.

In short: it's the chat-window loop with the copy-paste removed — and the safety is in how tightly you bound the "run commands" step.

2 · Install & first run essential

Codex ships as a CLI (plus IDE extension and a cloud version). Install it with whichever package manager you already use:

terminal# npm
npm install -g @openai/codex

# or Homebrew (macOS)
brew install --cask codex

# or the install script (macOS/Linux)
curl -fsSL https://chatgpt.com/codex/install.sh | sh

Then launch it and authenticate. The easiest path is to sign in with your ChatGPT account (Plus/Pro/Business/Edu/Enterprise plans include Codex access); alternatively you can use an OpenAI API key.

terminalcodex            # launches the interactive TUI; choose "Sign in with ChatGPT"
▶ How this works
  1. The install puts a codex binary on your PATH. Any one of the three methods works — pick the package manager you already maintain.
  2. Running codex with no arguments opens the interactive terminal UI, where you describe tasks in natural language and watch the agent work.
  3. On first run you authenticate once — "Sign in with ChatGPT" is simplest; an API key is the alternative for automation or non-ChatGPT accounts.

Try this: in a throwaway git repo, launch codex and ask it to "add a README with a one-line description of this project." Watch it propose the file before anything is written — that preview is the approval model doing its job.

3 · The agentic loop in practice intermediate

What makes the agent feel different from autocomplete is that it closes its own feedback loop. You give it a goal; it doesn't just emit a diff and stop. It reads the relevant files to understand context, proposes edits, runs the command that would tell it whether the edit worked (a test, a build, a linter), reads the output, and — if something failed — tries again with that new information. The same think-act-observe cycle you built by hand in the Agents chapter (Ch 4), now pointed at your codebase.

In practice you steer it with the goal and the guardrails, not the steps. A good task is outcome-shaped: "make test_auth.py pass", "refactor this module to remove the global", "add retry-with-backoff to the API client." The agent decides which files to touch and which commands to run; your job is to review its plan and bound its permissions — which is the next section.

It's the Ch 4 loop, in your terminalEverything you learned about the agent loop — think, act via a tool, observe the result, repeat, with a hard stop — applies here. The "tools" are reading files and running shell commands; the "stop" is the goal being met or your approval being withheld.

4 · Sandbox & approvals — the safety model advanced

This is the most important section on the page, because an agent that can run shell commands is an agent that can do damage. Codex gives you two independent dials — a sandbox that limits what the agent can do, and an approval policy that controls when it must ask you first. Set them deliberately; the defaults lean safe, and you widen only as you build trust.

The sandbox (sandbox_mode in config, or --sandbox) has three levels:

Sandbox modeWhat the agent can do
read-onlyRead files only — no edits, no commands that change state. Safest; good for "explain this codebase."
workspace-writeRead and write within the working directory, run commands — but no network/outside access by default. The everyday default.
danger-full-accessNo restrictions. Only for trusted, isolated environments (e.g. a disposable container).

The approval policy (approval_policy) controls when the agent pauses to ask permission before acting:

Approval policyWhen Codex asks you
untrustedAsks before anything not on a known-safe list — most cautious.
on-failureRuns, and asks for approval only if a command fails (e.g. to retry with more access).
on-requestThe agent decides when it needs to ask — it requests escalation for riskier steps.
neverNever asks — fully autonomous. Pair only with a tight sandbox or a disposable environment.

The two dials combine. The common beginner error is reaching for danger-full-access + never because the prompts are annoying — which hands an LLM unsupervised control of your machine. The professional default is workspace-write with approvals on, escalating a specific task to more access only when you've watched the agent behave. There's a --full-auto convenience flag for a low-friction (workspace-write + auto-approve) mode when you're iterating fast in a safe repo.

two independent dials — capability (sandbox) and when-it-asks (approvals) read-only workspace-write danger-full-access SANDBOX (can) untrusted on-failure / on-request never APPROVALS (asks) safe default: workspace-write + approvals on; widen only with trust Two dials, set independently. The sandbox caps what the agent can do; the approval policy caps when it acts without asking. Start green on both and move down only in a repo (or container) you can afford to lose.
🗺️ How to read this diagram
  • The left column is the sandbox — capability rising from green read-only to purple danger-full-access.
  • The right column is the approval policy — from green untrusted (asks the most) to purple never (fully autonomous).
  • They're independent: you choose a point on each. The dangerous corner is purple-plus-purple; the safe everyday setting is the middle/green zone.

In short: sandbox = what it can touch; approvals = when it asks. Default to the safe end of both and widen deliberately.

5 · Non-interactive runs for CI advanced

The interactive TUI is for you at your desk; CI needs the agent to run headless. Codex supports a non-interactive mode — codex exec "<task>" — that runs a task to completion without the TUI, which is what you wire into a pipeline (for example, "triage this failing build" or "open a PR that fixes the lint errors"). Because there's no human at the keyboard to answer prompts, the sandbox and approval settings become load-bearing: a headless run with never approvals must be paired with a tight sandbox or a disposable, isolated environment, or you've built an unsupervised agent with your credentials.

terminalcodex exec "fix the failing unit tests in this package"
Headless + full access = dangerIn CI there's no one to approve a risky step. Never combine approval_policy = "never" with sandbox_mode = "danger-full-access" on a runner that holds real secrets. Use workspace-write in an ephemeral container, scope the credentials to the minimum, and treat the agent like any other automated job with least privilege.

6 · config.toml & MCP servers professional

Codex reads a config file at ~/.codex/config.toml where you set defaults so you're not passing flags every time — the model, the sandbox mode, the approval policy, and any MCP servers you want the agent to use. MCP (the subject of the next chapter) lets the agent reach standardized external tools — a filesystem server, a GitHub server, a database server — each declared as an entry with a launch command and args.

~/.codex/config.toml# defaults so you don't pass flags every run
model = "gpt-5.5"
approval_policy = "on-request"
sandbox_mode = "workspace-write"

# an MCP server the agent can call (see ox4-mcp)
[mcp_servers.filesystem]
command = "npx"
args = ["-y", "@modelcontextprotocol/server-filesystem", "/path/to/project"]
▶ How this works
  1. The top keys set your defaults: which model, how cautious the approval_policy, and how capable the sandbox_mode — the same two dials from section 4, now persisted.
  2. Each [mcp_servers.<name>] table declares an MCP server the agent can use; command + args are how Codex launches it (here, a filesystem server over stdio).
  3. With this in place, codex starts already knowing your preferences and tool connections — no repeated flags.

Try this: set sandbox_mode = "read-only" in config and ask the agent to make a change — watch it explain what it would do but refuse to write. That's the config enforcing the dial.

7 · CLI agent vs API vs IDE tech-lead

Codex, the raw API, and an IDE assistant solve overlapping but different problems — knowing which to reach for is a leadership call. Reach for the CLI agent (Codex) when the task is a repo-level change with a verifiable outcome: fix the failing tests, do a mechanical refactor across many files, triage a build. Reach for the raw API (ox2) when you're building a product feature — the model is a component inside your app, not a tool at your terminal. Reach for an IDE assistant when you want inline, keystroke-level help while you drive.

For a team, the governance questions mirror section 4 at organizational scale: which repos may run an agent, in what sandbox, with whose credentials, and reviewed how. Treat an autonomous coding agent in CI like any other privileged automation — least privilege, auditable, and never holding more access than the task needs.

🪜 Practice ladder beginner → industry

  1. Beginner: install Codex and run it in read-only mode; ask it to explain an unfamiliar repo.
  2. Easy: in a throwaway repo, let it add a small file in workspace-write with approvals on; approve each step.
  3. Core: give it a failing test and have it iterate until green; watch the plan→edit→run→observe loop.
  4. Stretch: write a ~/.codex/config.toml with your preferred model, sandbox, and approval defaults.
  5. Hard: add an MCP filesystem server to the config and have the agent use it.
  6. Industry: design a safe codex exec CI job — ephemeral container, scoped creds, workspace-write, and an audit trail.

✓ Checkpoint — you can move on when you can…

  • Explain what an agentic coding CLI does and when to use one.
  • Install Codex and run a task interactively and with codex exec.
  • Set sandbox_mode and approval_policy deliberately and explain the safe default.
  • Configure ~/.codex/config.toml with defaults and an MCP server.

Knowledge check check yourself

✓ Knowledge check

What are Codex's sandbox modes, and what is the safe default pairing with the approval policy?

Show answer
The sandbox modes are read-only (read files only), workspace-write (read/write and run commands within the working directory), and danger-full-access (no restrictions). The safe everyday default is workspace-write with the approval policy on (e.g. on-request or untrusted), widening only with trust. Never pair danger-full-access with never approvals outside a disposable, isolated environment.
✓ Knowledge check

How do you run Codex non-interactively, and why do the sandbox/approval settings matter more there?

Show answer
Use codex exec "<task>" to run a task headless without the TUI — this is what you wire into CI. The settings matter more because there's no human to answer approval prompts: a headless run with never approvals must be paired with a tight sandbox (e.g. workspace-write) in an ephemeral, least-privilege environment, or it becomes an unsupervised agent holding your credentials.
© 2026 studybydoing.in · AI Engineering: Zero to Production · All rights reserved. · About · Privacy Policy · Terms · Contact
Educational content, provided as-is and without warranty. Code samples are examples — review, test, and adapt them before using in production. See the Terms of Use & Disclaimer. Use at your own risk.
© studybydoing.in