MEASUREMENT AT THE WIRE 2026-07-02

The Hidden Token Tax

49,907 tokens. That is what one Claude Code launch loads before you type a single word — measured at the wire, from the provider's own token accounting. Not an estimate. This post is what I found when I stopped trusting file scans and started reading the billing meter itself.

Cold start

49,907 tokens

Per Claude Code 2.1.198 launch, byte-identical across runs — the provider's own usage fields.

The over-count

29 on disk → 1 loaded

Moving every non-enabled plugin's skills off disk shifted cold start by 2 tokens.

The fix

4,959 → 319 tokens

One skill vault over MCP; same audit before and after; runtime-confirmed; reversible.

# The method, in one block [1] Loopback proxy between the CLI and the API → read the provider's usage fields [2] Change exactly one thing on disk → relaunch → read the meter again [3] Ask each runtime what IT says is loaded (inspectors, not file scans) [!] Every number below has a committed artifact and a re-run command _

01· Measured at the wire, not estimated

Token-counting tools model what an agent probably loads. I wanted the number the provider actually bills, so I built a small loopback reverse proxy, pointed ANTHROPIC_BASE_URL at it, and logged every response's native usage fields. Three cold launches of Claude Code 2.1.198 from an empty directory: 2,783 uncached + 47,124 cached prefix = 49,907 tokens, identical every time — system prompt, tool schemas, skills, memory, all loaded before the first user word.

Terminal proof: three Claude Code launches through the token proxy, each totaling 49,907 input tokens, with the raw usage JSON showing a 1-hour ephemeral cache
Proof 1 — the provider's own accounting, three launches, byte-identical. Also visible: Claude Code writes its prompt cache with a 1-hour TTL, not the API-default 5 minutes.

One caveat the data itself teaches: absolute cold-start size is a per-CLI-version snapshot — it moves as the CLI ships new tools and prompts. The mechanisms below, however, replicate — they are architecture facts, proven with deltas.

02· File scans lie; token deltas don't

My disk shows 29 Claude Code skills. The runtime loads exactly 1 (~78 tokens — the one enabled plugin). The proof is a controlled experiment, not an opinion: move every non-enabled plugin's skills off the disk, relaunch, read the meter — the cold-start context shifted by 2 tokens. Moving the 28 marketplace skill files themselves: exactly 0. Disabling the one enabled plugin: −75. Available ≠ installed ≠ enabled — and most audit tooling measures the disk, so it over-counts.

Codex is the counter-example that proves the design point: its entire 604-plugin marketplace catalog costs ~294 tokens at cold start (~0.5 per plugin — a compact index), and deleting its 5 built-in skills changed the prompt by exactly 0. Codex ships progressive disclosure by default.

Terminal proof: token-delta tables — Claude Code minus 75 tokens for disabling the enabled plugin, 0 for moving 28 marketplace skills; Codex 0 for removing built-in skills, minus 294 for the whole catalog
Proof 2 — one controlled filesystem change per probe, the CLI's own usage report before and after.

03· The cross-CLI picture

Across five CLIs — Claude Code, Codex, Gemini, Grok, Kimi — the on-disk skill-catalog tax on my machine topped out at 4,959 tokens per session (an upper bound: runtimes load subsets, which is exactly why the audit defers to each CLI's own inspector where one exists). The surprise: the biggest catalog isn't Claude's. Gemini carried ~3,000 tokens on disk while Claude's enabled catalog was 39.

Terminal proof: cross-CLI audit before the fix — claude 39, codex 397, grok 1482, gemini 3041 tokens, total 4959 per session
Proof 3 — the audit before the fix. Same tokenizer metric for every CLI; every total labeled an upper bound.

And the quieter tax: in the controlled harness, activating a set of 15 skills adds ~3,000 tokens to every session start whether your project rules are 100 tokens or 8,000 — and even never-invoked skills cost ~400 tokens of metadata. On my own machine, /doctor once dropped 30 skill descriptions because unused ones flooded the listing budget: the skills I needed lost context to the ones I didn't.

04· The fix, measured with the same instrument

The remedy is structural: consolidate every CLI's skills into one on-demand vault served over MCP — the catalog leaves the startup prompt entirely, and a skill's body enters context only when the model actually asks for it. Applied on my machine and re-audited with the same instrument: 4,959 → 319 tokens per session.

Terminal proof: same audit after the fix — 319 tokens total, only Grok's plugin registry remains
Proof 4 — same audit, after. The residual is Grok's plugin registry, which is never moved because that would break the plugin system.

Then I re-tested the runtimes themselves, because rule one of this work is that the audit is not the ground truth — the runtime is. Gemini's own inspector went from 9 skills loaded to 0. Grok's went from 38 to 23. And the re-test caught a survivor: Claude Code serves enabled plugins from a cache copy, so its one plugin skill was still loaded. I patched the tool to relocate the cache copy too and re-tested until the runtime confirmed it gone — 16 skills listed before, 15 after, the plugin's skill no longer among them. Stable cut once every quirk is counted: ~86%, every skill still available on demand, and the whole change reversible with one command.

05· Three rules from the bench

  • 1> Audit the runtime, not the disk. Probe with token deltas — the provider's usage numbers are the only ground truth.
  • 2> Curate skills per project. Every description you carry is context your code doesn't get.
  • 3> Keep the receipts. Screenshots, raw artifacts, and a re-run command for every number — a claim nobody can re-verify is marketing, not measurement.

06· The receipts

Everything above follows its own rule three. Behind each screenshot sits a dated raw artifact — the proxy's JSONL usage logs, the before/after audit JSON, the CLIs' own inspector outputs — plus a claim-to-evidence index mapping every number to its instrument and re-run command, and a cryptographic, tamper-evident logging chain for tool invocations.

The full write-up — methodology, raw measurement artifacts, and the verification chain — coming soon. If you want the short version in the meantime: measure the initial load of unused capabilities, then decide how much of it you actually need to pay for on every new session.