MEASUREMENT CLAUDE × CODEX 2026-07-02

The cold-start bill, model by model

Before you type a single word, your AI coding agent has already loaded a system prompt, tool schemas, skills and memory — and started billing for them. So I measured how much, per model, at the wire, from each provider's own token accounting. Four models. Four different bills. The spread runs from 11,662 to 49,907 tokens — a 4.3× range, all of it paid on every new session.

Leanest measured

11,662 tokens

Codex CLI, gpt-5.5 — a compact base runtime, progressive skill loading.

Heaviest measured

49,907 tokens

Claude Code 2.1.198, sonnet-5 — 2,783 fresh + 47,124 cached prefix, byte-identical across runs.

The honest lesson

It's a snapshot, not a constant

The number moves with the model, the CLI version, and even your flags. Measure your own.

# The method, in one block [1] Loopback proxy between the CLI and the API → read the provider's own usage fields [2] Launch each model from a fresh empty directory → the number IS the cold-start context [3] Tag every absolute with its model + CLI version + instrument — never collapse two [!] Every figure below has a committed raw artifact and a re-run command _

01· Four models, measured

Token-counting tools estimate what an agent probably loads. I wanted the number the provider actually counts, so I put a small loopback reverse proxy between the CLI and the API and logged every response's native usage fields. "Cold start" here means the full input context — system prompt, tool schemas, skills, memory — the runtime loads on the first request from a fresh empty directory, before any user work. Here is what each model actually loaded:

Model CLI / version Cold-start Instrument
codex gpt-5.5 Codex 0.142.5 11,662 CLI-relayed usage
claude haiku-4-5 Claude Code (June) 20,062 CLI usage report
claude fable-5 Claude Code 2.1.198 37,326 wire proxy
claude sonnet-5 Claude Code 2.1.198 49,907 wire proxy

n = 3 launches each (haiku n = 3, June); every cell byte-stable across its runs. Splits: codex 2,062 fresh + 9,600 cached · haiku 2,576 + 2,630 write + 14,856 read · sonnet 2,783 + 47,124 cached prefix.

Terminal proof: three Claude Code launches through the token proxy, each totaling 49,907 input tokens, with the raw usage JSON showing a 1-hour ephemeral cache
Proof — the sonnet-5 wire measurement: three launches, 49,907 tokens each, from Anthropic's own usage fields.

HONEST NOTE. These are four measured configurations, not a controlled model tournament. The two on the identical CLI (fable-5, sonnet-5) also differ in one launch flag, and the whole 12,581-token gap between them lives in the cached prefix — where tool schemas live — so I do not claim "the model caused the gap." Haiku is an older CLI version measured with the CLI's own usage report, kept here for historical context only. Isolating the pure per-model effect is a one-command controlled follow-up (swap only --model, hold flags constant). What every row above does prove: the cold-start number is real, repeatable, and specific to your model and version.

02· Why Codex starts 3× lighter

The gap between Codex's 11,662 and Claude's 37,326 isn't noise — it's architecture, and the controlled deltas prove it. I installed a fixed set of six synthetic skills into a project and relaunched, changing exactly one thing. Claude's cold-start context grew by 58 tokens. Codex's grew by 1,199 — the same six skills, a 20× difference in how eagerly each runtime pulls a project's skills into the startup prompt.

Claude — install 6 skills

+58 tokens

Heavier base, but a tiny per-skill catalog appetite — enablement-gated, mostly names.

Codex — install 6 skills

+1,199 tokens

Lean base, but it reaches for project skills far more eagerly when they're present.

So "leaner" cuts both ways: Codex ships the smaller base runtime, Claude ships the cheaper per-skill catalog. Which one wins for you depends on how many project skills you actually carry — which is exactly why a single published headline number is the wrong thing to optimize against. Measure your stack.

03· The turn where you actually use a skill costs double

Installing a skill is cheap-ish. Invoking one is not. I forced a single skill call and read the turn's full billed input. On both runtimes it landed at almost exactly twice a normal turn:

Runtime Normal turn Activation turn Multiple Verified
claude fable-5 38,095 75,801 1.99× 3/3
codex gpt-5.5 12,861 26,108 2.03× 3/3

The mechanism is visible in Codex's event stream: it activates a skill by shell-reading the SKILL.md file, then re-sends the context to the model — so the turn is billed roughly twice. That same progressive-disclosure design is why Codex's cold start is so lean: it doesn't pre-load skill bodies, it fetches them on demand. You pay for a skill when you use it, not before. Activation was verified by a magic string that exists only inside the fixture skill and had to appear in the model's final reply — 3 of 3 runs on each runtime.

04· The number moves — so pin it down

The same Claude Code that loads ~37–50k today measured 20,062 tokens on the June build (Haiku 4.5, from the CLI's own usage report). The runtime grew as it shipped new tools and prompts. That is the single most important caveat in this whole study: an absolute cold-start size is a snapshot of one (version, model) pair, not a law of nature. Cite it with its date and version or don't cite it.

The caching behind the number moves too. Claude Code writes its prompt cache with a 1-hour TTL (not the API-default five minutes) — but that stamp is a ceiling, not a guarantee: on a quiet machine the cache survived 30-minute idle gaps intact, yet under heavy same-account load it was fully evicted before 55 minutes. Codex's prefix cache is best-effort and block-level, with partial evictions as early as two minutes. Treat every cache TTL as an upper bound when you reason about cost.

The durable findings are the ones that survive a version bump: the deltas (§02, §03) and the mechanisms. The absolutes are receipts for a moment in time.

05· What to do with this

  • 1> Measure your own pair. Point a loopback proxy at the API, launch from an empty dir, read the usage field. Your model on your CLI version is the only number that bills you.
  • 2> Pick the model for the workload, knowing the floor. A leaner base (Codex) or a cheaper per-skill catalog (Claude) — the right trade depends on how many project skills you carry.
  • 3> Move the skill catalog off startup. Consolidating every CLI's skills into one on-demand vault served over MCP cut my cross-CLI catalog tax from 4,959 → 319 tokens per session — the full story is in The Hidden Token Tax.

06· The receipts

Every number here derives from a committed raw artifact — the proxy's JSONL usage logs, the live-study rows, the June cold-start harness — never hand-typed. The 35 measured launches are committed to an RFC-6962 Merkle tree + linear hash chain (root 771dd6a3…), so the dataset is tamper-evident: change one token and the root changes. Each figure maps to its instrument and a re-run command in a claim-to-evidence index.

# provenance cold-start absolutes → results/live-study-2026-07-02/rows.jsonl (fable-5, gpt-5.5) → results/token-proxy-claude-2026-07-02.jsonl (sonnet-5, 49,907) → results/aimeter-real-2026-06-20.json (haiku-4-5, June) catalog + activation → results/live-study-2026-07-02/report.json (deltas) claim → evidence → PROOFS.md (instrument + date + re-run command per number) tamper-evidence → RFC-6962 Merkle root 771dd6a3… over 35 measured rows

Not measured, not claimed: cold start for models that were only tokenizer-tested (Opus and others). "All models" here means the four whose cold start I actually put on the wire — no estimated rows padding the table.