01· Four models, measured
Token-counting tools estimate what an agent probably loads. I wanted the number the
provider actually counts, so I put a small loopback reverse proxy between the CLI and the API
and logged every response's native usage
fields. "Cold start" here means the full input context — system prompt, tool schemas, skills,
memory — the runtime loads on the first request from a fresh empty directory, before any user
work. Here is what each model actually loaded:
| Model | CLI / version | Cold-start | Instrument |
|---|---|---|---|
| codex gpt-5.5 | Codex 0.142.5 | 11,662 | CLI-relayed usage |
| claude haiku-4-5 | Claude Code (June) | 20,062 | CLI usage report |
| claude fable-5 | Claude Code 2.1.198 | 37,326 | wire proxy |
| claude sonnet-5 | Claude Code 2.1.198 | 49,907 | wire proxy |
n = 3 launches each (haiku n = 3, June); every cell byte-stable across its runs. Splits: codex 2,062 fresh + 9,600 cached · haiku 2,576 + 2,630 write + 14,856 read · sonnet 2,783 + 47,124 cached prefix.
usage fields.
HONEST NOTE. These are four measured
configurations, not a controlled model tournament. The two on the identical CLI
(fable-5, sonnet-5) also differ in one launch flag, and the whole 12,581-token gap between
them lives in the cached prefix — where tool schemas live — so I do not claim "the
model caused the gap." Haiku is an older CLI version measured with the CLI's own usage
report, kept here for historical context only. Isolating the pure per-model effect is a
one-command controlled follow-up (swap only --model, hold
flags constant). What every row above does prove: the cold-start number is real,
repeatable, and specific to your model and version.
02· Why Codex starts 3× lighter
The gap between Codex's 11,662 and Claude's 37,326 isn't noise — it's architecture, and the controlled deltas prove it. I installed a fixed set of six synthetic skills into a project and relaunched, changing exactly one thing. Claude's cold-start context grew by 58 tokens. Codex's grew by 1,199 — the same six skills, a 20× difference in how eagerly each runtime pulls a project's skills into the startup prompt.
+58 tokens
Heavier base, but a tiny per-skill catalog appetite — enablement-gated, mostly names.
+1,199 tokens
Lean base, but it reaches for project skills far more eagerly when they're present.
So "leaner" cuts both ways: Codex ships the smaller base runtime, Claude ships the cheaper per-skill catalog. Which one wins for you depends on how many project skills you actually carry — which is exactly why a single published headline number is the wrong thing to optimize against. Measure your stack.
03· The turn where you actually use a skill costs double
Installing a skill is cheap-ish. Invoking one is not. I forced a single skill call and read the turn's full billed input. On both runtimes it landed at almost exactly twice a normal turn:
| Runtime | Normal turn | Activation turn | Multiple | Verified |
|---|---|---|---|---|
| claude fable-5 | 38,095 | 75,801 | 1.99× | 3/3 |
| codex gpt-5.5 | 12,861 | 26,108 | 2.03× | 3/3 |
The mechanism is visible in Codex's event stream: it activates a skill by
shell-reading the SKILL.md file,
then re-sends the context to the model — so the turn is billed roughly twice. That same
progressive-disclosure design is why Codex's cold start is so lean: it doesn't
pre-load skill bodies, it fetches them on demand. You pay for a skill when you use it, not
before. Activation was verified by a magic string that exists only inside the fixture skill
and had to appear in the model's final reply — 3 of 3 runs on each runtime.
04· The number moves — so pin it down
The same Claude Code that loads ~37–50k today measured 20,062 tokens on the June build (Haiku 4.5, from the CLI's own usage report). The runtime grew as it shipped new tools and prompts. That is the single most important caveat in this whole study: an absolute cold-start size is a snapshot of one (version, model) pair, not a law of nature. Cite it with its date and version or don't cite it.
The caching behind the number moves too. Claude Code writes its prompt cache with a 1-hour TTL (not the API-default five minutes) — but that stamp is a ceiling, not a guarantee: on a quiet machine the cache survived 30-minute idle gaps intact, yet under heavy same-account load it was fully evicted before 55 minutes. Codex's prefix cache is best-effort and block-level, with partial evictions as early as two minutes. Treat every cache TTL as an upper bound when you reason about cost.
The durable findings are the ones that survive a version bump: the deltas (§02, §03) and the mechanisms. The absolutes are receipts for a moment in time.
05· What to do with this
- 1> Measure your own pair. Point a loopback proxy at the API, launch from an empty dir, read the
usagefield. Your model on your CLI version is the only number that bills you. - 2> Pick the model for the workload, knowing the floor. A leaner base (Codex) or a cheaper per-skill catalog (Claude) — the right trade depends on how many project skills you carry.
- 3> Move the skill catalog off startup. Consolidating every CLI's skills into one on-demand vault served over MCP cut my cross-CLI catalog tax from 4,959 → 319 tokens per session — the full story is in The Hidden Token Tax.
06· The receipts
Every number here derives from a committed raw artifact — the proxy's JSONL usage logs, the
live-study rows, the June cold-start harness — never hand-typed. The 35 measured launches are
committed to an RFC-6962 Merkle tree + linear hash chain
(root 771dd6a3…), so the dataset is
tamper-evident: change one token and the root changes. Each figure maps to its instrument and
a re-run command in a claim-to-evidence index.
Not measured, not claimed: cold start for models that were only tokenizer-tested (Opus and others). "All models" here means the four whose cold start I actually put on the wire — no estimated rows padding the table.