Tools and configuration to make an AI agent (Claude Code) spend far fewer tokens — without degrading the quality of its answers.
Born from a multi-agent audit of real Claude Code usage (50+ subagents over 47 transcripts). The finding that drives everything:
~88% of the tokens in a 3D-modeling session go into READING IMAGES (renders and reference photos at full resolution, ~110k tokens each). It's not the text, the config, or the verbosity — it's vision.
goldentokens attacks that spend and the recurring improvisation patterns around it, with a handful of small, deterministic CLIs the agent calls instead of rewriting throwaway code.
git clone https://github.com/VortexJer/Goldentoken.git && cd Goldentoken && pip install -r requirements.txt && python goldentoken.pyThat clones the repo, installs the dependencies, and launches the installer, which opens an
interactive menu to pick a risk level (safe / medium / dangerous). Press ESC to accept the
safe default. Restart Claude Code afterwards. To pick the level non-interactively, append a flag:
python goldentoken.py --safe|--medium|--dangerous. Uninstall anytime with python goldentoken.py --uninstall.
| Tool | What it does | Why it saves |
|---|---|---|
img 🔴 |
Cheap image inspection: view (downscale + dedup by content hash), crop, grid (ruler in original coords), montage, sheet, diff (metrics), convert. |
Reading a ≤768px derivative costs ~2-5k tok instead of ~110k. The project's biggest saving. |
jsonq |
JSON/JSONL queries by path (parts[].name), --where, --pick, schema. Replaces jq (not installed). |
Avoids rewriting json.load by hand (measured: 240+ times). |
editpre |
Preflight for an Edit: validates old_string and returns the exact on-disk bytes (fixes cp1252 quotes, dashes, CRLF). |
Avoids the Edit-fails → re-read whole file → retry cycle. |
runq |
Runs a command and prunes its verbose output always keeping the errors + last N lines. | Avoids dumping hundreds of build/pytest/npm lines into context. |
sightreport |
Reads the report.json of the *sight tools (auto-discovers the latest); compact WARN/FAIL-only table. |
Avoids rewriting json.load of the report (measured: 83 times in 19 sessions). |
peek |
Skeleton of a large file (~5% of tokens) so you range-read only what you need. | Avoids Read of the whole file (62% of reads were full-file). |
pathcheck |
Verifies paths before a cd/ls/cat that assumes they exist. | Avoids the 'No such file or directory' failure. |
envkv |
Read/check/add keys in .env without duplicating or corrupting (idempotent). |
Replaces improvised grep/append hacks. |
toktrack |
Token telemetry from a Claude Code transcript (measures, changes nothing). | — |
refprep |
Prepares a reference photo (silhouette/mask/measure) for modeling. | Replaces ~24 throwaway PIL scripts. |
waitfor |
Waits for a local port/server to be ready. | Replaces improvised urlopen loops. |
skillreload |
Reloads a *sight skill after editing: kill + pip --force-reinstall + verify import, in one command. |
Removes a 3-command ritual per reload. |
ci-await |
Waits for a GitHub Actions run (optional publish) in one command. | Replaces the manual gh run dance. |
ship |
commit/push + their opposites uncommit/unpush, dry-run by default. |
Replaces the repetitive git ritual, safely. |
aiapi |
Model/pricing catalog of any AI provider via its official API, not HTML scraping. | Replaces scraping thousands of tokens of docs HTML. |
webshot |
Web screenshot (playwright) in one command, ready to read cheaply with img. |
Replaces throwaway screenshot scripts. |
Key design of img view --hash: it dedups by content hash, not by filename. A render regenerated with the same name is NOT a duplicate; it only skips byte-identical content reads.
python bin/img.py view render.png --max 512 # cheap thumbnail
python bin/img.py view render.png --hash # 'UNCHANGED' if it didn't change
python bin/img.py crop render.png --box 300,300,300,350 --zoom 2
python bin/img.py grid render.png --step 100 --major 500 --label
python bin/img.py diff a.png b.png --align silhouette
python bin/jsonq.py report.json 'validation.errors' --len
python bin/editpre.py file.js --old 'const x = "hello"'
python bin/runq.py --last 8 -- pytest -q
python bin/aiapi.py list openrouter --grep claude
python bin/ship.py commit -m "msg" --new-branch feature/x # dry-run; add --yes to runAll accept both Windows-style and Git-Bash-style paths (/c/Users/...).
pip install pillow numpy # dependencies (only `img` needs them)
python goldentoken.py # interactive level menu (ESC/ignore -> safe)
python goldentoken.py --safe # -s low-risk hooks only
python goldentoken.py --medium # -m up to MEDIUM risk (includes -s)
python goldentoken.py --dangerous # -d everything, incl. HIGH risk (includes -m)
python goldentoken.py --stats # -sts estimated cumulative savings
python goldentoken.py --uninstall # remove everything(If the repo is on your PATH, you can type goldentoken -d thanks to goldentoken.cmd.)
Running with no flag opens a menu (arrows + Enter) to pick a level; if ignored or you press ESC, it installs safe by default. Each install also:
- installs a real lock (
capability-guard): the AI CANNOT run operations above the level (e.g.ship pushon safe/medium is denied, not just discouraged); - installs the
goldentokensskill that teaches the AI the catalog, stamped with the active level; - records estimated savings (see
--stats).
Each component activates only if its risk ≤ the chosen level. Re-running with a different level reconfigures (removes what's above, adds what's below) without stacking:
| component | tier | what it does | risk |
|---|---|---|---|
| capability-guard | 🟢 safe | the real lock: denies operations above the installed level | none |
| output-cap | 🟢 safe | warns about huge outputs (cuts nothing) | 0 degradation |
| shell-guard | 🟢 safe | denies high-confidence Windows shell footguns (with the fix) | low |
| read-guard | 🟠 medium | blocks re-reading a byte-identical file (R3) | medium |
| concise (output-style) | 🟠 medium | trims prose; opt-in via /output-style concise (R4) |
medium |
| img-redirect | 🔴 dangerous | auto-downscale of images: the biggest saving but can lose fine detail (R1) | high |
Honest note: the big token saving (images, ~88% of real spend) lives in the img-redirect
hook, which is the only dangerous one. Under --safe/--medium the hooks barely compress
images automatically — but the img CLI is available to call by hand at any level. The installer
uses 8.3 short paths without quotes (immune to spaces), backs up your settings.json, writes
a manifest, and does not overwrite your other hooks (e.g. globalcontext). Restart Claude Code
after installing.
install.py / uninstall.py are kept as aliases (install.py = medium level).
tests/bench_savings.py measures token savings locally, without spending AI tokens (official
w·h/750 formula for images; byte ratio for text). Summary:
| tool | avg saving |
|---|---|
| read-guard / jsonq / sightreport / editpre | 99–100% (slice/dedup of something large) |
| runq | 95% (prunes verbose logs, keeps errors) |
| img view 768px | 90% (77% small render → 96% phone photo) |
| peek | 59% (skeleton; keeps all signatures) |
Detail and honest method in tests/SAVINGS.md. Quality (same conclusion as without the tool) is
covered separately in RISK-TRACKING.md (deterministic parity + subagent stress tests).
- 16 CLIs — working and tested (151/151 in
tests/test_all.py) - 5 hooks:
capability-guard,shell-guard,img-redirect,read-guard,output-cap— wired - risk-level installer/uninstaller (
goldentoken.py) with interactive menu +--stats— tested - measured savings (
tests/SAVINGS.md) and quality parity (RISK-TRACKING.md)
Dropped from the audit catalog (user's decision): refspec, tabular-enrich,
solidsight validate (the last one under the "solidsight is off-limits" rule).
MIT (to be added).