For Claude Code · measure → trim → re-measure

Measure the tax · trim it · keep the same quality

Claude Code Token Tuner — Cut Your Token Bill Without Losing Quality

Claude bills you for context size × turns. A bloated CLAUDE.md, auto-loaded rule files you forgot were on, and a few silent cache-busters quietly re-bill you every turn. The Token Tuner makes the tax visible, then shows you exactly how to trim it — same model, same skills, same quality.

Instant download · offline measuring tool · ~2-minute setup · no API key for the core

Runs offline No API key (core) Real audit inside ~2-min setup
Product cover

The problem

You can't trim what you can't see

Most people never measure their always-on token tax — so they pay for the same bytes on every turn of every session, forever.

▌ Flying blind

  • A bloated CLAUDE.md re-loads on every single turn
  • Auto-loaded rule files you forgot were even on
  • Silent cache-busters re-bill your whole history at full price
  • No number in front of you — so nothing ever gets fixed

▌ Token Tuner

  • weigh_context.py counts your always-on tax — offline, no key
  • A two-track Trim Playbook: paste-one-prompt beginner + advanced manual
  • Names the silent cache-busters most token counters never show you
  • Run it, trim, run it again, watch the number drop

What it does

Measure, trim, and dodge the silent spikes

Built from a real audit: a live setup measured down from ~13,900 to ~7,300 always-on tokens (−47%). Token counts are chars/4 estimates for before/after, not exact billing.

01 · MEASURE

See the tax

weigh_context.py counts your always-on token tax. No API key, no tokens spent, runs offline. Run it, trim, run it again, watch the number drop.

02 · TRIM

Two tracks

BEGINNER: paste one prompt, do three free things — permissions.deny read-rules, a concise output style, one-task-per-session. ADVANCED: the deeper wins most people miss.

03 · SPIKES

The hidden cache-busters

A mid-session model switch, the opusplan trap, and enabling fast mode mid-session all invalidate your prompt cache — so the next turn re-reads history without the cached prefix.

04 · WINS

The new levers

Fork subagents that share the parent's cache prefix (up to ~90% less input on parallel work), a hook that shrinks a 10,000-line test log before it hits context, and Batch API at 50% off.

What's in the box

Everything to measure and trim

⚠ Requirements — please read before buying

  • Claude Code (this is a Claude Code optimization kit)
  • macOS or Linux — Windows works via WSL
  • Python 3 (already installed on Mac/Linux)
  • No API keys for the core (one optional advanced tactic, Batch API, needs API billing)

Stop paying for the same bytes twice

Make the token tax visible — then cut it

Same model, same skills, same quality. You just stop re-paying for context you didn't need loaded. Unzip, install, read your number, trim, re-measure.

Instant download · 30-day Gumroad refund policy applies

Setting up Claude Code from scratch? Start with the Claude Code Setup Playbook — $29.

See all my Claude Code tools → expressive446.gumroad.com