compression with a quality contract

Core Concepts

Three ideas explain why Distil's numbers are real and why the approach is different: decision-equivalence as the correct quality target, a statistical certificate that enforces it, and cache-aware cost modeling that makes aggressive compression financially correct.

How Distil works — a cache-aware compression pipeline with a quality contract AGENTDISTIL PROXY · localhostPROVIDER Claude Code · Codex Gemini · any SDK re-sends full context every turn compress → cache-pin → forward Tier-0 — lossless verbatim transforms (always on)Tier-0lossless Tier-1 — reversible decision-aware digestTier-1 digestreversible Causal / counterfactual pruning — ablation-provencausal pruneablation-proven Salience protection — model-free needle-keepingsaliencemodel-free Cache-aware placement — pin the stable prefix at the 0.1× cache-read rate cache-aware placement — stable prefix billed at the 0.1× cache-read rate RestoreStore — originals kept locally, never sent to the modelRestoreStoreoriginals kept locally · 0 tokens on wirenever sent to the model distil_expand — the model recovers any digested detail on demanddistil_expandmodel recovers any digested detailtransparent re-query · PAYG Anthropic · OpenAI Google Gemini prompt cache + LLM response → transparent expand loop → agent (it just gets the right answer) THE QUALITY CONTRACT — every operating point is certified decision-equivalent Decision-Equivalence Certificate — the conformal guarantee Decision-Equivalence Certificate conformal risk control (LTT/CRC) · distribution-free P(decision-change ≤ α) ≥ 1 − δ Shadow mode — live decision-equivalence on real traffic Shadow mode — live samples real traffic, runs it both ways rolling decision-change rate · streaming-aware Learned keep-model — the self-improving, never-regressing flywheel Learned keep-model — flywheel every expand is a label; only ever gets more conservative — never-regressing

Interactive — select any block (click, or Tab then Enter) to open the concept or technique behind it.


1. Decision-equivalence, not byte-equivalence

Most context compressors frame the problem as a string transformation: shrink the context, compare the output byte-by-byte, report overlap as accuracy. That framing has a fatal flaw — byte-equivalence and high compression are information-theoretically in tension. A claim of "87% compression, 100% accuracy" under byte-equivalence is self-contradicting; the two numbers cannot describe the same operation.

Decision-equivalence is the correct target. An agent only needs to take the same actions and produce the same outputs on the compressed context as it would on the original. "Same decision" means:

This is measurable, certifiable, and compatible with aggressive compression. Context blocks that are causally inert — they never change a decision when removed — are provably free to drop, regardless of how many bytes they contain. Distil's salience protection identifies the load-bearing frontier (model-free, using a causal ablation pass rather than a scorer model) and ensures those blocks are never touched, regardless of compression depth.

The reframe in one sentence: "100% accuracy" becomes a statistical non-inferiority guarantee on the agent's next action, not a diff of strings. On a live claude-opus-4-8 grading run, Distil hit 83.2% token savings at a measured 0% decision-change rate (2026-07-05, distil 1.10.1 vs llmlingua 0.2.2 and headroom-ai 0.27.0) — with the ≤5%-at-95% bound certified offline by the conformal risk control (DERC).
Honest scope (we measured this). Next-action equivalence is a proxy. On a real multi-turn task — SWE-bench Verified, end-to-end (E7) — the proxy certificate does not transfer to task success once compression is aggressive (pass@1 52%→16% at trunc@500). The contract is right; the lesson is that a certificate earned on a single-turn proxy must be re-earned on the real task. Distil's defensible value: cache-aware savings inside a proxy-certified gate, plus a reversible recovery tier. E8 (500-instance long-horizon ReAct agent, 6 conditions, full SWE-bench Verified) is that re-earned test. The relevance-gated tier (E) achieves 36.8% pass@1 — the highest of any compressor tested — and is not distinguishable from full context (difference −2.4 pp, 95% CI [−5.7, +0.9]; McNemar p=0.19), though the interval's lower end is past a 5 pp margin, so non-inferiority at 5 points is not shown. It is a research configuration; the shipped path's task outcome is on SWE-bench Lite (cost-neutral, within about 2 points). It also beats Headroom (+4.2 pp, p=0.035). LLMLingua-2 removes a nearly identical share of context (52% vs 53%) yet reaches only 2.4% pass@1 — 34 points lower, isolating what is kept as the deciding factor. Lossy truncation collapses to 5.6%. Distil does not claim cheapest; Headroom is cheaper. The claim is certified accuracy with a guarantee. E10 extends this to a trajectory-level certificate (distribution-free, same LTT/HB engine as E2): with 95% confidence, the gated compressor costs a solvable task on ≤11.4% of exchangeable runs, proven out-of-sample across 1000 calibration/test splits.

Two recently shipped techniques apply decision-equivalence directly to the coding-agent read → edit → re-read loop. Cache-delta context coding (distil/cachedelta.py) replaces an exact or near-duplicate re-read with a back-reference or reference + diff. It is decision-equivalent because the prior version of the file is already present earlier in the cached conversation — prior version + diff carries the same information. It also enforces cache-monotonicity: only the volatile suffix is mutated, so every prompt-cache hit that existed before the turn still exists after it. The deepest layer, AST-structural delta (distil/astdelta.py), fingerprints Python definitions with ast.dump (invariant to whitespace, comments, and import order), so a reformat-only re-read is a zero-change reference rather than a spurious diff. For a full treatment see Techniques → Cache-delta and AST-structural delta.

Three fidelity tiers, precisely. Lossless — the model sees content byte-identical in-context (Tier-0 / --verbatim). Reversible — content is digested but byte-recoverable on demand via the local store / distil_expand (the default on PAYG; like a zip you unpack only when needed). Lossy — dropped irrecoverably (every other tool). All three Distil tiers are certified decision-equivalent; Headroom also ships a local store + retrieve (uncertified, dev-facing); only Distil certifies the reversible tier and makes recovery agent-facing — an injected tool the model calls mid-task. (--lossless-only, alias --safe, restricts to Tier-0 verbatim — no lossy shaping, no tool injection, and no Tier-1 stubs; it folds directly into --verbatim. Use for subscription / OAuth sessions.) Two rules sit outside the tier ladder and bind all of it: a block whose bytes are a verbatim slice of a file is never digested at all, so an Edit quoting it still matches (provenance is read from the shell command, not just the tool name); and no keep budget is shared between blocks, so a large untrusted block cannot starve a trusted one — see the threat model and the degradation curve for what each rung of the ladder actually costs.

2. The quality contract

A compression strategy in Distil does not ship unless a pre-registered statistical gate certifies it. Distil uses two complementary certification methods, both grounded in conformal prediction:

The Student-t tail used in TOST is computed from a hand-rolled regularised incomplete beta function — zero dependencies on scipy or numpy. Both gates run across all 7 corpus domains.

How the TOST gate works

  1. Pre-register the indifference margin and significance level.
  2. For each turn, run the agent on the original context and on the compressed context; record whether the decision matched.
  3. Compute paired differences (compressed score − baseline score) across all turns.
  4. Run the lower one-sided test. Reject H₀ at α → the strategy is certified non-inferior.
$ distil certify --strategy distil
decision-equivalence match rate: 100.0%
TOST non-inferiority (margin=0.02, alpha=0.05): mean diff=+0.000, p=0.0000

VERDICT: PASS  (certified non-inferior)

$ distil certify --strategy aggressive
decision-equivalence match rate: 0.0%
TOST non-inferiority (margin=0.02, alpha=0.05): mean diff=−1.000, p=1.0000

VERDICT: FAIL  (NOT certified — would degrade quality)

A strategy that cannot pass the gate simply does not ship. You can watch the gate reject a quality-degrading strategy yourself — it is not a configuration option.


The trajectory-level certificate

DERC certifies the step — that the next action is unchanged. Distil's own E7 experiment, and the 2024–26 agent-compression literature (e.g. arXiv 2412.17483, 2510.00615), shows why that is not the whole story: a statistically valid per-step certificate can pass while end-to-end task success collapses, because per-step fidelity metrics systematically overpredict multi-step success. Certifying next-action equivalence and implying task fidelity is a category error.

distil certify-trajectories certifies the invariant that actually transfers: run your eval suite twice — full context and compressed — on the same tasks, and get a distribution-free bound on how many solvable tasks compression may cost you, using Conformal Risk Control and Learn-Then-Test (Angelopoulos et al., arXiv 2208.02814) on the matched full/compressed outcomes.

$ distil certify-trajectories outcomes.jsonl --alpha 0.05 --delta 0.05
# each line: {"task_id": "...", "full_success": true, "compressed_success": true}
distil trajectory-risk certificate

  matched trajectories : 200
  observed degradation : 0.50%
  risk bound (1-δ)     : 3.20%
  certified (α=0.05, δ=0.05) : True

  With confidence 95%, compression degrades at most 5.0% of tasks the full
  context would have solved (observed 0.50% over 200 matched trajectories).
Honest scope — the assumptions ship inside the certificate, not as fine print.
  • Exchangeability. The bound holds for future trajectories exchangeable with the calibration set. Agent workloads drift; pair the certificate with the built-in anytime-valid drift monitor and recalibrate when it alarms — a static certificate silently loses coverage under shift. The live alarm bets against the same budget the certificate is written at (distil.conformal.BUDGET_ALPHA), and on the proxy it acts: a breach holds compression at lossless-only, across restarts, until distil calibrate then distil reset --drift-guard (opt out: DISTIL_NO_DRIFT_GUARD=1).
  • The loss is the task outcome you measured. compressed_success should come from a real end-to-end evaluation (test suite passed, task rubric met), not a proxy metric — proxies are exactly the failure this certificate exists to fix.
  • No adaptivity. Trajectories used to tune the compressor must not be reused as calibration data — that breaks exchangeability a second way. Keep a held-out calibration stream.
Below min_n=20 matched trajectories the command refuses to certify rather than overclaim — a certificate from a handful of samples is noise dressed up as a guarantee.

Matched failures also feed the outcome-guided compression policy (always on): when a task the full context solved fails under compression, the content classes digested in that trajectory are suspects, and classes whose digestion keeps co-occurring with regressions get protected — kept byte-exact, automatically. Evidence is stored as content-free signatures (json:l, error:m), never content, so it composes with the expand flywheel without adding new state to reason about. Like that flywheel, it is never-regressing by construction: it can only make Distil more conservative, never less accurate.

See CLI Reference → distil certify-trajectories for the full flag reference.


3. Cache misses — not size — dominate agent cost

This is the insight most context compressors miss — and where Distil's biggest wins come from.

In a multi-turn agent loop, the growing context is re-sent to the model at every turn. With prompt caching enabled:

The dominant cost lever is therefore not how many tokens you send — it is how many of those tokens hit the cache. A naive compressor that rewrites the context every turn drops the longest common prefix to zero, losing the 10× discount on every previously-cached block.

Naive recompression sends fewer tokens yet costs more than no compression at all, because it rewrites the cached prefix every turn and pays full input price on tokens that would otherwise cost 0.1×.

Distil's cache-aware engine models this explicitly. It keeps the prefix byte-stable across turns via schema canonicalization and volatile-field extraction, compresses only the volatile tail, and proves via simulation that the naive approach is strictly dominated.

Confirmed on the live path from 1.45, and worth stating because it was not always true there: the provider's reported cache_read shows the prefix surviving every turn (a live A/B whose raw output was not committed, so no percentage is quoted here). Earlier versions reproduced the naive failure this section describes.

There is one moment the rule relaxes. A provider's cache entry lives for a fixed time after the last request that used it; a turn that arrives after it has expired pays to re-write the whole prefix no matter what is sent. Cold-point recompression uses exactly that turn to shrink older tool output behind recoverable stubs, then holds the smaller prefix byte-stable from there on. It acts only when expiry is certain from distil's own clock, and it is not yet measured live — see the cache contract, clause (h).

The numbers, measured

33%
Cost cut, cache-aware lossless · ops/sre trajectory
−11%
Naive compress + cache (costs more than baseline)
10×
Cache-read discount vs. fresh input (Anthropic pricing)
Cache-aware savings vs naive compression

The 33% figure is the conservative lossless, single-trajectory result. Across a varied 64-trajectory corpus, the certified lossless ceiling is ~58%. The distil-causal mode reaches ~81% on that offline corpus. And it holds up where it counts — under live model grading: in the live head-to-head against the real installed packages (a realistic, decision-determined corpus, graded by claude-opus-4-8), Distil certified 83.2% savings at a 0% decision-change rate, while LLMLingua-2 reached only 53.1% savings and flipped 1-in-8 decisions, and Headroom 0.27.0 achieved 39.7% at 0% — 2026-07-05, distil 1.10.1 vs llmlingua 0.2.2 and headroom-ai 0.27.0. Against headroom-ai 0.37.0, re-run 2026-09-04: distil-causal takes 52.9% tokens / 58.7% dollars at 100% decision-equivalence (PASS) vs Headroom 1.7% / 2.0% / 81% (FAIL) on the warm corpus gate, and on the read→edit→re-read codebench workload Headroom takes 35.6% tokens for +4.9% dollars.

For the compression mechanics, see Techniques. To price any trajectory yourself, see CLI Reference → distil savings.


Token Economics · Module 1 of 3 ~4 min read

Fundamentals — what a token is

In this module
  • What a token is — Byte-Pair Encoding, and the anatomy of one.
  • Why JSON, logs and code cost roughly 2× the tokens of prose.
  • Why tokenizers aren't interchangeable — and can shift under you.
  • Where the tokens actually go in a single agent request.

What a token actually is

A token is not a character and not a word — it's a subword unit, the atomic chunk a model reads and predicts. Modern models build their vocabulary with Byte-Pair Encoding (BPE): start from individual bytes and greedily merge the most frequent adjacent pair until you hit a target vocabulary size. Frequent words collapse into one token; rare words fall back to a handful of subword pieces. There is no <unknown> — every possible string is representable, in any language.

RAW STRING (a tool result) {"status":"ok","rows":2} BPE tokenizer 13 TOKENS — STRUCTURE COSTS EXTRA {" status ":" ok "," rows ": 2 } Only ok and 2 are signal. The braces, quotes and keys are pure overhead — and you pay for every one. CHARACTERS PER TOKEN — DENSITY VARIES BY CONTENT ~4.0 English prose 100 tokens ≈ 75 words. The efficient case. ~2.5 Source code camelCase, operators & indentation split hard. ~2.0 JSON / logs Punctuation-dense. The same bytes cost ~2× the tokens.

Three consequences fall out of how BPE works, and they drive the whole course:

Where the tokens go — anatomy of one request

Here's where most cost surprises come from. A single agent turn is not "the user's question." The API is stateless, so every turn re-sends the entire working set. A typical request, before the model reads a word of your actual task, already carries:

One request to the model — token budget Widths ≈ proportions commonly seen in tool-heavy agents. The query you care about is the sliver on the right. System prompt Tool schemas Message history Tool results persona, rules every tool's JSON schema — often the biggest block all prior turns, resent verbatim ← your query A 5-server tool setup can spend ~55,000 tokens before the task is read.

The biggest, most-overlooked sources of waste, with the hardest public numbers we have:

Key takeaways
  • A token is a BPE subword unit; structured text (JSON, code, logs) costs ~2× the tokens of prose — and agents live in that expensive regime.
  • Count tokens with the target model's tokenizer; they're neither portable across families nor fixed within one.
  • Most of an agent request is scaffolding (system prompt, tool schemas, history, stale tool results) — not your task. That gap is where compression lives.

Token Economics · Module 2 of 3 ~7 min read

Cost & Compression

In this module
  • Why multi-turn cost grows with the square of the conversation.
  • How prompt caching rewrites the math — and what silently breaks it.
  • The full family tree of compression, mapped by risk vs. savings.
  • Why a trained compressor beats a heuristic — and how it fails.

Module 1 showed where the tokens pile up inside a single request. Now we follow them across a whole conversation — where the real money is.

The multi-turn tax: why cost is quadratic

Single-turn cost is linear and boring. Multi-turn is where bills explode — and the shape is quadratic, not linear. Because each turn re-bills all the tokens before it, the cumulative cost of an n-turn loop is the sum 1 + 2 + … + n = n(n+1)/2 = O(n²). A four-turn conversation costs 4 + 3 + 2 + 1 = 10× the first turn's tokens — not 4×.

Cumulative billed input across a conversation Each turn re-sends everything before it. The area under the curve is what you pay. 0 turn 1turn 2turn 3 turn 4turn 5turn 6 Naïve — fresh input every turn (1.0×) Cached prefix — repeated tokens read at 0.1× N TURNS n(n+1) 2 = O(n²)

This is the foundation of Distil's approach — see Concepts → cache misses dominate cost for the measured numbers behind it.

Try it: the cost of a loop

Real arithmetic — no marketing. This computes the n(n+1)/2 input cost from your inputs, using live catalog pricing and the documented cache multipliers (0.1× read / 1.25× write). The "with Distil" line applies a flat lossless reduction to fresh tokens; the default 26.8% is Distil's measured corpus aggregate (see Research). Assumptions are shown so you can check the math.

$0
Naïve — no caching, no compression
$0
With prompt caching (prefix read @ 0.1×)
$0
With caching + Distil lossless (−26.8%)

Assumptions: the starting values are arbitrary calculator defaults, not a measured session; input-token cost only (output excluded for clarity); prefix is cacheable and byte-stable; per-turn input = prefix + accumulated history; cached line writes the prefix once at 1.25× then reads it at 0.10× while the growing tail stays fresh; Distil line additionally removes 26.8% of fresh tokens losslessly. Illustrative — your numbers depend on your traffic.

Compression, the whole family tree

"Compression" is an umbrella over a dozen different ideas, each trading a different amount of quality risk for savings. The most common mistake in this space is conflating them — input-token compression (fewer tokens enter the model), KV-cache compression (the prompt is read in full; the retained attention memory shrinks), and cost reduction (same tokens, cheaper compute) are three different axes. Here is the honest map.

Risk vs. savings — where each family sits more savings → ↑ more quality risk certified-safe band — provably no decision change Lossless minify / dedup / RLE Retrieval (RAG) & semantic cache Extractive pruning (drop low-info tokens) Abstractive summary / rolling memory Learned keep/drop classifier Soft-prompt / gist vectors KV-cache eviction (H2O, SnapKV) — memory, not tokens
Lossless · mechanical

Make the same bytes cost fewer tokens

Whitespace and JSON normalization, key shortening, de-duplication, run-length collapse, compact serialization. Risk: essentially zero — information is preserved exactly. The ceiling is modest, but it is free savings you should always take first.

Extractive prompt compression

Drop the least-informative tokens

Score each token's informativeness (self-information / perplexity) and prune the bottom. Systems like Selective Context and the LLMLingua family reach 2–20× on prose. Risk: real — a low-surprisal but load-bearing token (a negation, a constraint) can be dropped, and output is often human-unreadable.

Abstractive · summarization

Rewrite the context shorter

Summarize retrieved docs or older turns into a compact form (RECOMP; rolling-summary agent memory; OS-style tiered memory like MemGPT). High ratios. Risk: lossy — specific facts, numbers and entities can be distorted or dropped, and errors compound across turns.

Retrieval as compression

Keep knowledge outside the window

RAG and semantic caching hold the corpus in an index and pull only the top-k per query — lossy compression of a whole knowledge base into what fits. Risk: retriever-bounded — a missed chunk becomes a hallucination; a too-loose semantic-cache hit returns a stale answer.

KV-cache compression

Shrink the attention memory, not the input

H2O, Scissorhands, SnapKV, StreamingLLM keep only "heavy-hitter" or recent keys/values in GPU memory. Different axis: the model still reads every input token, so this cuts memory/latency, not your input bill. Eviction is permanent — a wrongly-dropped key can't return.

Soft-prompt / gist

Compress context into learned vectors

Gisting, AutoCompressors and ICAE fold a long prompt into a few learned "summary" vectors (up to ~26× reported). Risk: model-specific & trained — the artifact only works for the model it was trained on, and reconstruction is lossy.

And two cost-reduction cousins that aren't token compression at all: speculative decoding (a small draft model proposes tokens the big model verifies in parallel — provably identical output, 2–3× faster) and model cascades / routing (FrugalGPT, RouteLLM — try a cheap model first, escalate only when needed). They lower cost without removing tokens, and compose cleanly with everything above.

Why training a small model helps

The most effective compressors stopped using fixed heuristics and started learning. The root idea is knowledge distillation (Hinton et al., 2015): train a small "student" to imitate a large "teacher," using the teacher's soft probabilities — which carry far richer signal than a one-hot label.

Applied to compression, the move is to reframe it as per-token keep/drop classification. LLMLingua-2 is the clean example: take a small bidirectional encoder (XLM-RoBERTa, ~355M params), and train it on labels distilled from GPT-4 — prompt GPT-4 to compress text by word-removal only, then mark each original token "keep" or "drop." A trained classifier beats a perplexity heuristic for three concrete reasons:

The risk that comes with learning. A trained compressor internalizes its training distribution and fails quietly. The documented failure modes are real: out-of-distribution regression; over-compression dropping exactly the load-bearing tokens (high-information tokens are often the most important, yet get pruned at the same rate); and task-dependent fragility (classification and code degrade far harder than summarization). Past a point, aggressive compression flips from cost-saving to cost-increasing as garbled outputs trigger retries. This is why a learned keep-model must be gated, never trusted blindly — the subject of Module 3.
Key takeaways
  • Multi-turn cost is O(n²) because every turn re-bills all prior tokens — a 4-turn loop is 10×, not 4×.
  • Prompt caching is the biggest lever (read @ 0.10× vs fresh @ 1.0×), but any byte change in the prefix throws the cache away.
  • Compression is a family of techniques on three different axes; pick by your risk tolerance, lossless-first.
  • A trained keep/drop model beats heuristics — but fails quietly, so it must be gated by proof.

Token Economics · Module 3 of 3 ~5 min read

Distil & Proof

In this module
  • Why "it still looks fine" is not proof — and what is.
  • Non-inferiority testing, borrowed from clinical trials.
  • How Distil applies every technique in certified, risk-graded tiers.
  • A practical playbook, and where the field is heading.

Proving it didn't hurt

Every technique in Module 2 can degrade quality invisibly — the output still looks reasonable. So the load-bearing question isn't "how much did you save?" It's "how do you know the model still makes the same decision?" Three things to internalize:

How Distil applies it — in tiers

Distil's design choice is to apply the safe techniques aggressively, the risky ones only under proof, and to refuse anything it can't certify. That's the tier ladder: each rung saves more and risks more, and only ships if it clears the gate above it.

⛓ Quality contract — non-inferiority (TOST) gate No tier ships unless the agent's decisions are certified statistically not-worse. Fails → fall back a rung. must pass ↓ 0 Tier 0 — Lossless JSON minify · collapse exact-duplicate runs. Byte-recoverable. Always on. risk: none applied to every request 1 Tier 1 — Reversible digest Verbose tool output replaced by a compact digest + handle the agent can expand on demand. Reconstructable. risk: bounded reversible by construction ★ Certified — Causal / counterfactual pruning Ablation finds context the decision doesn't depend on; cache-aware & keep-model guided. Kept only if it passes the gate. risk: gated to zero certified per run

The pieces that make the ladder work — each documented in depth on Techniques (which also carries the full tier reference table):

# See the certified frontier and what's safe to take on your own corpus
distil eval
distil bench            # lossless + causal savings, with the non-inferiority verdict
distil proxy --port 8788 # drop-in; records genuine savings as real traffic flows
distil leaderboard      # your real cumulative savings (local, private)

A practical playbook

Independent of any tool, this is the order of operations that saves the most for the least risk:

Do first — free wins

  • Stabilize your prefix. Put the system prompt and tool schemas first and keep them byte-identical; move anything volatile (timestamps, IDs) after the last cache breakpoint. This alone can 10× the discount on repeated tokens.
  • Turn on prompt caching and verify hits via the cache-token counts.
  • Minify structured payloads losslessly before they enter context.
  • Prune dead tool definitions — every unused schema is paid on every turn.

Then — measured wins

  • Digest verbose tool output reversibly; let the agent expand on demand instead of carrying full dumps forever.
  • Compact or summarize stale history — but treat summaries as lossy and keep recent turns verbatim.
  • Prune causally, not by surprisal — remove what the decision doesn't depend on.
  • Certify every aggressive step with a non-inferiority test; if it can't pass, don't ship it.
  • Measure genuine savings on real traffic, not a synthetic corpus.

Where the field is heading

You've finished the course. The throughline of every advance in this field is one move: find the tokens the outcome doesn't depend on, and stop paying for them — without changing the outcome. Distil makes that last clause a proof, applied in tiers, measured on your real traffic. Next: Install & Quickstart, then go deeper in Concepts and Techniques.

Further reading

Primary sources behind this course — verified, for the curious: