Open source · for Claude Code, Codex, Gemini CLI and more

The compressor that measures itself
— and every other one.

Distil trims the tool output, logs and history your agent re-sends every turn, without breaking your prompt cache, keeps every byte it trims so the agent can get it back, and reports whether that saved you money or changed what the agent did. It runs any other compressor through the same harness, by the same rules.

What the measurements say today: on short SWE-bench tasks no compressor is shown to save money once a shared-prompt-cache confound is removed — distil is at parity, RTK is −6.5% per task [−13.1%, +0.1%], not significant (it read −20.7% before the fix). A 7-task long-session pilot is inconclusive. An offline simulator that reproduces the provider's bill within 2.5% on held-out runs finds lossless policies move cost by under 0.2%, and lossy savings hinge on whether the agent takes extra steps — which the next live A/B measures. On the maintainer's own Claude Code traffic (13,191 requests, September 2026) distil removed an estimated 8.3–9.2% of the bill, assuming the agent behaved the same (how). Scoreboard → · Why RTK looked cheaper →

uv tool install distil-llm && distil setup

Or curl -LsSf https://dshakes.github.io/distil/install.sh | sh · brew install dshakes/tap/distil · Windows: powershell -ExecutionPolicy ByPass -c "irm https://dshakes.github.io/distil/install.ps1 | iex"

~/your-project — distil
$ distil setup            # detects your agent + billing, wires everything
$ distil wrap -- claude    # or let setup make it always-on
$ distil savings           # what it saved you, from your own traffic
$ distil doctor            # if something looks off
distil wrap -- claude also keeps Claude Code's MCP tool search switched on, which Claude Code otherwise turns off behind a proxy, so unused connectors can stay deferred (verified live on 1.54.0: tools_deferred 4–5 on every request, tool payload 9,733 tokens, no request failures; ADR 0013).

Why trust the number

  • It doesn't break your prompt cache. Distil never rewrites bytes the provider still has cached; older context changes only once that cache has already expired — an enforced invariant. The cache contract →
  • Every compressed byte is recoverable. Folded content stays in a local store; the agent calls distil_expand to get the exact original back mid-task. How →
  • It's measured on your own bill. Savings are calibrated against the usage your provider actually bills, not estimated from a benchmark. How it's measured →

New to this? How it works in plain English → · The proof, in depth ↓

Works with
  • Claude Code
  • Codex
  • Gemini CLI
  • aider
  • Cursor
  • CrewAI
  • LangChain
  • LangGraph
  • LiteLLM
  • Agno
  • AWS Strands
  • Vercel AI SDK
  • MCP

…and anything else that lets you set a base URL — distil is a proxy, so there is no per-framework package to install or keep in sync. Two carry caveats worth reading before you rely on them: Cursor routes only its agent/chat panel through a custom base URL (tab-completion and ⌘K stay on its own backend), and CrewAI needs the LLM passed to planning_llm and function_calling_llm too, not just the agents. Codex's preset is unverified: it speaks the right wire shape, but whether OPENAI_BASE_URL reaches codex's request path at all is still open — distil wrap --list says so. Full integration matrix →

The proof, in depth

Distil is the compressor that tells you when it is hurting you — and measures any other compressor, or your provider's own compaction, the same way.

On SWE-bench Lite the shipped path is cost-neutral and within about 2 points on task success — not statistically distinguishable at n=300. Two runs after the shell-search fix: 204/299 vs 210/299 plain (−2.0 pts, 95% CI −5.5..+1.5, cost $7.18 vs $7.18) and 213/298 vs 218/298 (−1.7 pts, 95% CI −5.1..+1.7, cost $8.44 vs $8.52). Non-inferiority at the pre-registered 5-point margin is not shown, and a 7-task long-session pilot is inconclusive. Reports and method → · every compressor, same harness →

Four things are proven rather than promised, and each has its own page: a cache contract (compression may never rewrite bytes the provider already cached), a COMA-class adversarial gate (with the two cases that don't come back clean published), a degradation curve across every rung of the dial, and an exact-quote guarantee so an Edit still matches the file the agent read.

Honest scope: the certificate covers decision-equivalence on a trajectory corpus, not end-to-end task success — on real SWE-bench Verified the proxy does not transfer once compression gets aggressive, and we publish that. Research configurations that did better there are not what ships — full statement, CIs and ablations →

~9%
Off a real bill
maintainer's metered traffic · Sep 2026 · net of distil's own spend
37/40
Decisions changed by Anthropic's context clearing
measured with certify-provider · keep=0 · n=40
~1000×
Faster compression
0.026 ms/turn · no model in the path
0
Runtime dependencies
stdlib-only · 700+ tests · pip · docker · pyz
The difference

Most compressors ask for trust.
Distil measures — and says when it hurts.

✗ Trust · every other compressor
  • Drops tokens and hopes the agent still decides the same
  • “Looks fine” on a demo — no proof on your traffic
  • You discover the regression in production
trust me · no receipt
✓ Prove · distil
  • Certifies decision-equivalence before it ever serves
  • A shadow gate re-checks live decisions with paired replays — a self-agreement control absorbs sampling noise
  • Gives back what it folds — byte-exact, from a content-addressed store, through a tool the agent can call mid-task
  • Tested against hostile input, not just hard input — and the two cases that fail are on the page
  • Fails closed — if quality would drop, it doesn't compress
signed certificate · live, per mode (2026-09-15): digest −5.3 pp vs self-agreement, over budget · lossless-only ≈0
In plain English

How it works — in plain English#

No jargon. Here's the whole idea in three pictures.

Don't re-send what's already remembered

Every turn, your agent re-sends the entire conversation — and you pay for all of it, again. The model can cheaply "remember" the parts that haven't changed, so Distil keeps those parts perfectly stable and only touches the new bits. Like not re-reading someone the whole book to add one sentence.

Drop only what doesn't matter

Most of that context never changes what the agent decides to do next. Distil finds those dead-weight parts with a simple test — remove it, replay, did the decision change? If no, it was safe to cut. Keep the clues, drop the filler.

Prove it — don't just trust it

Here's the part nobody else does: Distil measures that your agent still makes the same choice on the shrunk context. If it can't prove that, it sends everything, untouched. A warranty on the compression, not a promise.

And nothing is truly thrown away — trimmed detail is filed nearby, and the agent can pull back the exact original the moment it needs it. That's why Distil can compress hard and stay safe. The gentle, longer walkthrough →

The architecture

Under the hood — a cost-optimized cache hierarchy with a contract#

Most compressors optimize bytes. In an agent loop the money is somewhere else.

Architecture: agent traffic flows through the Distil proxy — drift guard, cold-point planner, salience, Tier-0 lossless and Tier-1 digest, prefix replay — to the provider, with a local restore store and recovery loop, a sealed receipt chain, and the decision-equivalence contract whose live drift alarm can hold the proxy at lossless-only

Two highest-leverage techniques — and why they win

TECHNIQUE #1

Cache-aware compression

You re-send the growing context every turn. With prompt caching a cache read is ~10× cheaper than fresh input — so the dominant cost is cache misses, not context size. Distil keeps the prefix byte-stable and compresses only the volatile tail. Naive recompression sends fewer tokens yet costs more, because it rewrites the cached prefix every turn.

Since 1.45 the live adapter keeps the cached prefix byte-stable, and the provider's own cache_read confirms the prefix survives every turn (a live A/B whose raw output was not committed). Before 1.45 the live adapter had the failure it names above — in that same A/B it cost 2× direct. See CACHE.md.

TECHNIQUE #2

Causal / counterfactual pruning

The eval engine isn't a ruler — it's a discovery engine. Remove a context block, replay, did any decision change? Blocks that never change a decision are provably free to drop: speculative retrievals, stale history. The measurement produces the compression policy.

TECHNIQUE #3

Exact-quote provenance

A coding agent's Edit is a literal match against bytes it read earlier. Digest that read and the decision is unchanged — the edit is simply skipped, and the run ends with the agent reporting success and the file untouched. Distil reads provenance from the shell command (cat, head, sed -n), not just the tool name, and refuses to digest a verbatim slice of a file.

Measured content-free over 2,489 local sessions: shell reads are 33.6% of tool-result mass, against 10.7% matched by a tool-name rule. Byte-exact quote loss falls 39.3% → 16.2%. It costs savings — that price is published rather than hidden.

TECHNIQUE #4

An estimator allowed to say no

Shadow mode replays a sampled request three times — twice on the original context, once on the compressed one — and reports the paired difference 1{A=B} − 1{A=A′} with a bootstrap 95% CI, unclipped. A ratio clamped at zero can never report harm; a difference can. One reporting floor (50 A/B + 30 A/A) gates every surface, and below it each one prints below reporting floor instead of a number. It reports per compression mode, because pooling hides the one that hurts: on the maintainer's 2026-09-15 sample, digest agreed 102/206 times against a self-agreement of 113/206 (−5.3 pp, over the budget) while lossless-only read 109/192 against 109/193 (≈0). The drift guard now holds digest at lossless-only wherever its recent harm bound is over budget.

TECHNIQUE #5

A grader that is itself measured

A certificate is only as good as the model that graded it. The live certifier was changed under a pre-registered, paired migration eval on real traces, and the served path is graded separately from the certified one. Model migration →

The reframe that makes "100%" real. Byte-equivalence and high compression are information-theoretically in tension. Decision-equivalence is the right target: the agent takes the same actions and produces the same outputs whether or not the context was compressed. That is measurable and certifiable — "100%" becomes a statistical non-inferiority guarantee on outcomes, not a string diff.
Receipts, not vibes

The proof — a quality contract, not an estimate#

We didn't just benchmark ourselves. We ran the real LLMLingua-2 and Headroom packages through the same gate, graded live by the real model.

Live head-to-head: Distil vs LLMLingua-2 vs Headroom
On the decision-equivalence proxy, Distil is the most aggressive, fully decision-equivalent, and lowest-latency. On a synthetic 120-turn offline corpus, certified 83.2% token savings at a 0% decision-change rate (≤5% at 95% confidence) — while LLMLingua-2 cuts 53% but flips 1-in-8 decisions. See the full live benchmark →
Dated: 2026-07-05, distil 1.10.1 vs llmlingua 0.2.2 and headroom-ai 0.27.0. Headroom has moved since: re-run 2026-09-04 against headroom-ai 0.37.0, distil-causal takes 52.9% tokens / 58.7% dollars at 100% decision-equivalence (PASS) against Headroom’s 1.7% / 2.0% / 81% (FAIL) on the warm corpus gate, and on the read→edit→re-read codebench workload Headroom takes 35.6% tokens for +4.9% dollars. The “2.1× less aggressive” ratio is from 0.27.0 and does not survive the re-run. Fresh head-to-head →
Proxy caveat: this is next-action equivalence on a (partly synthetically decision-determined) corpus — not end-to-end task success. On real SWE-bench Verified (E7) aggressive compression drops pass@1 52%→16%; the proxy certificate does not transfer. Research configurations (E8–E14) did better on SWE-bench Verified; they are not the shipped path, and the shipped path's task outcome is the SWE-bench Lite result at the top of this page.

Underneath it: a strategy ships only if a pre-registered non-inferiority test certifies it. Lossless passes; quality-degrading compression is rejected.

And after it ships, the same budget keeps watching. The certificate, the live drift alarm and the proof ledger read one pinned risk budget (decision change ≤5% at 95% confidence). When shadow mode's paired verdicts prove live traffic has crossed it, the proxy stops compressing lossily from the next request, across restarts, and writes the trip to the receipt chain. Every wrap exit says so in plain words until you recalibrate. An alarm that only prints at exit is a log line; this one acts.

$ distil savings --pricing claude-opus-4-8
strategy                               $ / run   vs baseline  cache hits
------------------------------------------------------------------------
baseline (no cache, no compress)       0.01524          0.0%           0
cache only                             0.01115         26.8%       1,028
naive compress + cache                 0.01696        -11.3%           0   ← busts cache
distil (cache-aware lossless)          0.01023         32.8%       1,028

$ distil certify --strategy distil
decision-equivalence match rate: 100.0%
VERDICT: PASS  (certified non-inferior)

$ distil certify --strategy aggressive
decision-equivalence match rate: 0.0%
VERDICT: FAIL  (would degrade quality — blocked)
Bar chart of per-run cost: baseline (no cache, no compress) $0.01524; cache only $0.01115 (-27%); naive recompression busts the prefix cache and costs more, $0.01696; distil's cache-aware lossless approach $0.01023 (-33%), the lowest of the four
Adoption · live

Small numbers, shown anyway#

These sit here rather than beside the proof tiles above, because they answer a different question. The tiles above are what distil does, and they are large. These are how many people have opted in to being counted, and they are small — a project that publishes a calibrated decision-change rate does not get to hide its install count. Every figure is recomputed nightly by public code from a public git branch, and consenting is not contributing: an idle opt-in is never counted as a saver.

Coverage

Measured across 9 domains — the same gate, not one example#

A strategy isn't trustworthy because it works once. distil bench certifies every domain — ops, coding, support, research, data, devops, finance — or it doesn't ship.

Bar chart of lossless savings across 9 domains, all passing the decision-equivalence gate: ops/sre 32.8%, coding 25.5%, support 32.6%, research 25.7%, data-analysis 18.1%, devops 22.8%, finance 24.9%, web-research 89.8%, agent-worklog 35.3%; aggregate 48.4% cheaper reversibly across 7,080 prunable tokens
Risk grading

Risk-graded tiers#

Apply the provably-safe ones everywhere; gate the rest on evidence.

Tier 0

Provably lossless

Reconstructable transforms — JSON minify, reversible run-length encoding. Always on.

Tier 1

Reversible digest

Decision-aware digest + a handle; the full original stays local and re-expands on demand.

Certified

Lossy, but gated

Pruning & summarization allowed only at ratios the non-inferiority gate certifies.

The cost stack

Capabilities — every layer of the cost stack#

CapabilityWhat it doesLoss profile
Priced cache-aware engineModels the multi-turn loop and proves naive recompression busts the cache—
Schema canonicalizationRecursively key-sorts JSON/tool payloads so the prefix is byte-stablelossless
Volatile-field extractionLifts dates/UUIDs/JWTs out of the cached prefix so it stops churninglossless · reversible
Reversible digest + handlesOriginals stay local and re-expand on demandreversible
Reject-if-bigger invariantNever emits a block larger than its originalsafety
Adversarial validation gatedistil validate — drives the real compressor against hostile inputs and asserts reversibility, recency-exactness, fail-open, and content-free telemetry holdCI gate
Causal / counterfactual pruningAblation discovers context that never changes a decisioncertified
TOST non-inferiority gate (DERC)Per-step: a strategy ships only if it passes the quality contractthe moat
Trajectory-level certificateTask-level: distil certify-trajectories bounds end-to-end degradation on matched full/compressed runsthe differentiator
Savings ledger + leaderboardLocal-first, privacy-preserving cumulative savingsopt-in
Content-type keep policyPins each kind's load-bearing lines — a log's pass/fail verdict, a traceback's frames, a diff's hunk headers1.15
Query-aware salienceSees the agent's intent in the same request and pins the line it's asking about — lexically (a grep hit, a SHA), and now semantically: a learned model keeps the line that answers the query with no shared word ("retry limit" → max_attempts = 5)1.16
Columnar foldRe-encodes JSON record arrays into a compact self-describing table — real token savings even on a flat-rate planlossless
Session dissectdistil dissect — per-session deep-dive with an anomaly list that catches silent failures automatically1.15
Savings discoverydistil discover — what is still costing you across recent sessions, ranked, each with its derivation and the command to act on it; median and p10/p90 printed beside the best session, and silent where it cannot measurereport
Drift guardWhen live paired verdicts prove decision change has crossed the one risk budget, every proxy on the machine serves lossless-only until distil reset --drift-guard — an alarm that acts, not a log linesafety · rc soak
Cold-point recompressionShrinks older tool output only on a turn the provider's cache has certainly expired for, then holds those bytes stable — clause (h); not yet measured livereversible · rc soak
Sealed receipt chainA content-free, hash-chained receipt per request, sealed into segments with Merkle-root checkpoints: hand an auditor one receipt and its inclusion proof, not the whole log (distil receipts --prove)audit
Opaque-block passthroughAnthropic compaction/thinking blocks and OpenAI reasoning/compaction items reach the provider byte-identical, pinned by contract tests, and are censused rather than hiddenuntouched
Tool definitions left aloneThe tools array is never rewritten (measured: too little to gain, ADR 0012); distil wrap -- claude keeps Claude Code's own MCP tool search on instead — verified live on 1.54.0 (tools_deferred 4–5, tool payload 9,733 tokens)untouched
Drop-in

Works with every SDK — one proxy, no code change#

Point any base_url-honoring client at the proxy — Python, TypeScript, any language — and get cache-aware lossless compression. See integrations →

Diagram: five client SDKs (Anthropic, OpenAI, Vercel AI, LangChain, LiteLLM) all point base_url at the Distil proxy on localhost:8788, which does cache-aware compression, lossless reversible transforms, and auth-mode gating (or runs in-process via wrap(client)), then forwards to Anthropic, OpenAI-compatible, or Google Gemini APIs with no code changes
Get started

Install — pick a format#

The stdlib-only core makes the packaging clean: zero runtime deps, corpus bundled.

uvx — zero install, recommended. Runs straight from PyPI, nothing to clean up.

uvx --from distil-llm distil bench

Run straight from PyPI — nothing installed. Needs uv.

pipx install distil-llm
distil certify

Permanent, isolated venv — PEP 668-safe. brew install pipx if needed.

brew install dshakes/tap/distil
distil bench

Isolated keg — doesn't touch your system Python.

docker build -t distil .
docker run distil bench

Container image — reproducible, zero host deps.

make pyz
python dist/distil.pyz bench

Single file — copy distil.pyz anywhere with Python 3.9+.

Already installed? Do this next → distil onboard detects your agent + billing, wires the savings status line, and gives you a guided tour tailored to your setup — then distil default makes it the default for every session (distil offboard reverses that).

What install actually does: every door drops the proxy into an isolated environment — uv/pipx venv, a Homebrew keg, or a container — never your system Python (pip install is blocked by PEP 668 on modern macOS/Linux, and that's fine). distil onboard wires your agent and the savings status line; distil default makes it the default and distil offboard fully reverses it — nothing is touched that isn't put back. Node/other stacks just point their SDK baseURL at distil proxy. Verify anytime: distil bench → expect GATE: PASS.
Drop-in, no call-site change:
from distil.adapters.anthropic import wrap
client = wrap(anthropic.Anthropic())  # compresses + keeps the cache warm
No overselling

What we won't pretend#

A headline like "87% reduction, 100% accuracy" is unverified marketing until you see the eval. Published quality numbers in this space are typically measured at far lower compression than the headline workload ratios — so the two figures don't describe the same run. Distil's answer isn't a bigger number; it's a gate: the accuracy claim and the compression are measured on the same trajectories across 9 domains, and a strategy that can't pass non-inferiority doesn't ship. The default tokenizer is an offline heuristic (ratios robust, dollars approximate); --tokenizer anthropic and --runner anthropic make it billing-grade against the real model.
Prove it on your own traffic

Compress the context.
Keep the decisions.

One proxy in front of any agent. A signed decision-equivalence certificate on the way out — or it doesn't compress. No API key to try it.