compression with a quality contract
Live head-to-head

The Benchmark

Every compression method shrinks tokens — the only number that matters is how much it saves without changing what the agent decides, once prompt caching is priced in. This head-to-head runs every technique through the same decision-equivalence gate and the same cache-aware cost model: the winner is computed, not assumed.

The result, in one line. On a realistic, decision-determined agent corpus graded live by claude-opus-4-8, Distil is certified at 83.2% token savings with a 0% decision-change rate (≤5% guaranteed at 95% confidence) — while running ~1,000× faster than the nearest tool. The real competitors, run live through the same gate on 2026-07-05, distil 1.10.1 vs llmlingua 0.2.2 and headroom-ai 0.27.0: LLMLingua-2 cuts 53% but flips 1-in-8 decisions (fails); Headroom 0.27.0 is decision-safe but only 40%. Headroom has moved since: re-run 2026-09-04 against headroom-ai 0.37.0, distil-causal takes 52.9% tokens / 58.7% dollars at 100% decision-equivalence (PASS) against Headroom’s 1.7% / 2.0% / 81% (FAIL) on the warm corpus gate, and on the read→edit→re-read codebench workload Headroom takes 35.6% tokens for +4.9% dollars. The “2.1× less aggressive” ratio is from 0.27.0 and does not survive the re-run. See the live head-to-head and full report.
Read this before quoting 83.2%. Every number on this page is decision-equivalence — the agent's next action on a trajectory corpus (partly synthetically decision-determined). That is a proxy, not end-to-end task success. In our real SWE-bench Verified end-to-end test (E7), aggressive compression drops actual pass@1 from 52% to 16% and the proxy certificate does not transfer. E8 (500-instance long-horizon ReAct agent, 6 conditions, full SWE-bench Verified) is the end-to-end positive counterpart: distil's relevance-gated tier achieves 36.8% pass@1, the highest of any compressor tested, is non-inferior to full context (−2.4 pp, 95% CI [−5.7, +0.9]; McNemar p=0.19), and beats Headroom by +4.2 pp (p=0.035). LLMLingua-2 removes a nearly identical fraction of context (52%) yet reaches only 2.4% pass@1. Lossy truncation craters to 5.6%. These are research configurations, not the shipped path; the shipped path’s task outcome is on SWE-bench Lite: cost-neutral and within about 2 points (reports). Distil does not claim cheapest — Headroom is cheaper — it claims certified accuracy with a guarantee. The full methodology — including why compression ratio alone is meaningless and how the A/A nondeterminism baseline works — is in docs/EVALUATION.md.
Token savings are not dollar savings. This page measures how much context each method removes without changing the agent's next decision. Whether that lowers the bill once the prompt cache and the agent's own behaviour are in is measured end to end on the scoreboard: on short SWE-bench tasks no compressor is shown to save money once a shared-cache confound in that run is removed (RTK −6.5% per task, [−13.1%, +0.1%], not significant; distil at parity), and a 7-task long-session pilot is inconclusive. Why →

Live head-to-head — real packages, real model

Not reference implementations: the actual installed packages (llmlingua 0.2.2, headroom-ai 0.27.0), each invoked the way that gives it its best fair result, all graded live by claude-opus-4-8 (majority-of-3) on the same realistic corpus (5 domains, 120 turns, 4.5–6.5 KB/turn). Savings are measured identically for everyone; the decision is the agent's actual next {action, target}.

Bar chart: Distil achieves 83.2% token savings with 0.0% decision-change (certified, 0.026 ms/turn); LLMLingua-2 reaches 53.1% savings but 12.5% decision-change (fails the gate); Headroom reaches 39.7% savings with 0.0% decision-change (certified, 26 ms/turn); full figures in the table below

MethodToken savingsLive decision-changeCertifies ≤5%@95%?Latency/turn
Distil (causal-prune + lossless) 83.2%0.0% ✔ certified — leader0.026 ms
LLMLingua-2 (llmlingua, real) 53.1%12.5% ✘ fails gate~1,480 ms
Headroom (headroom-ai, real) 39.7%0.0% ✔ certified26 ms

Live, claude-opus-4-8, majority-of-3, n=120, α=0.05/δ=0.05. Distil is the only method that is simultaneously the most aggressive, fully decision-equivalent, and lowest-latency. Headroom 0.27.0 is genuinely decision-safe here (it preserved the exact target ID every turn) but 2.1× less aggressive and loads a ModernBERT scorer — both facts are specific to 0.27.0 and to this 2026-07-05 run; see the 2026-09-04 re-run against headroom-ai 0.37.0. LLMLingua-2 is aggressive but decision-unaware — it drops/garbles the load-bearing ID on 1-in-8 turns, so the gate disqualifies it. Reproduce: python benchmarks/derc_live_compare.py; raw output of the 2026-07-05 run is committed at docs/paper/results/. Full methodology & caveats →

Why RTK isn't in the table. RTK (rtk-py) is a command-output proxy — it compresses the output of specific wrapped commands (git, ls, psql, aws…) by stripping known boilerplate, and exposes no raw-text/stdin mode. It physically cannot compress arbitrary agent context, so it's a different layer of the stack, not a contender on this axis. We attempted it for real; the adapter reports the mismatch rather than fabricating a number.

The certified frontier (live)

Every level of Distil's compression ladder, graded live — and the cliff where blind compression falls off it:

ladder level        decision-change   savings   certifies @ ≤5%/95%
--------------------------------------------------------------------
byte-exact                  0.0%        2.6%     ✔ yes
lossless (tier-0/1)         0.0%       47.9%     ✔ yes
causal-prune                0.0%       83.2%     ✔ yes
prune + lossless            0.0%       83.2%     ✔ yes  ← operating point
truncate@400 (blind)        0.0%       61.6%     ✔ yes
truncate@200 (blind)      100.0%       73.9%     ✘ no — drops the directive
--------------------------------------------------------------------
byte-exact = 0% confirms decisions are DETERMINED by context (the test is valid);
blind truncate@200 cuts more raw tokens than lossless yet flips EVERY decision.
It's not how many tokens you cut — it's which ones.

How we certified — and why it's credible

The loss

Decision-equivalence, not bytes

Per turn, the loss is 1 iff the agent's {action, target} flips versus the uncompressed context — graded by the same live model, majority-of-3 to strip the model's own run-to-run noise. We certify the decision, not a string diff.

The precondition

Determinism, verified

byte-exact = 0% proves the decision is determined by context — so any change at higher compression is the compressor's fault, not model ambiguity. (Our first corpus failed this honestly at ~50%; we fixed the corpus, not the math.)

The statistics

Distribution-free, finite-sample

Learn-Then-Test with Hoeffding–Bentkus p-values gives P(R(λ̂) ≤ α) ≥ 1−δ — no distributional assumptions, valid at finite n (arXiv:2110.01052 / 2208.02814). hb_p = 0.0058 < δ: certified with margin.

Why believe it

Built to be falsified

Every level shown (no cherry-pick); the certificate refuses when data is thin (honest conservatism); competitors run the same gate; fully reproducible with pinned versions. Caveats stated plainly: decision-determined synthetic corpus, exchangeability, marginal-not-per-prompt.

Coding-agent benchmark — cache-delta on read→edit→reread

A second, messages-level benchmark targets the coding-agent hot path (benchmarks/codebench.py: 20 sessions / 320 turns of read → edit → re-read, re-run 2026-09-04), scoring cache-aware real dollars against the real installed packages: headroom-ai 0.37.0 ([ml,code] extras), llmlingua 0.2.2, distil 1.51.1. Each method is a pure, deterministic, cache-monotonic function of the cumulative conversation — the same standard distil holds itself to. Raw output committed at benchmarks/results/2026-09-04/.

Cache-delta: read, edit, re-read — distil sends a reference + diff
methodtoken savings$ savings (cache-aware)latency / turnfidelity
LLMLingua-2 (llmlingua 0.2.2, real)56.7%57.2%173.7 mslossy
distil-verbatim + cache-delta42.7%46.3%58.2 msreversible
distil + cache-delta34.9%43.8%59.4 msreversible
Headroom (headroom-ai 0.37.0, real, default)35.6%4.9%6.9 mslossy
distil-verbatim (Tier-0 only)18.5%3.7%0.5 msreversible
distil (PAYG digest)0.0%*0.0%*1.0 msreversible

*By design, not a regression: this workload is read → edit → re-read, and once the agent has quoted a Read/Grep/Glob result byte-exact, that provenance is exempted from digest for the rest of the session — the exact-quote guarantee an Edit depends on. See the 1.49.0 changelog entry. Headroom's ms/turn above pays a one-time ModernBERT weight-load cost on this process's first invocation; a warm re-run (same token/$ percentages — see codebench-warm-headroom-2026-09-04.out) measures 2.2 ms/turn. distil has no model weights to load, so its own numbers don't move cold or warm.

Honest reading, 2026-09-04. Once digested reads are exempt from re-digesting, the reversible floor on this workload comes from cache-delta, not the bare digest: distil + cache-delta holds 43.8% cache-aware $ savings, fully reversible. The bare PAYG digest correctly measures 0.0% here — the 1.49.0 provenance-exemption fix protects the byte-exact quote an Edit needs, at the cost of the digest's savings on a workload that is nothing but re-reads. LLMLingua-2 still posts the highest raw compression (56.7% tokens, 57.2% dollars) — we don't hide it — but it's lossy and, even with the digest now doing far less work here, ~175× slower than it. Headroom's own number is a real improvement over its June 2026-06-23 run below: token savings flipped from a net cost increase to a net decrease (4.9% cheaper) on a newer, non-ML code path — though its default read-lifecycle compression still leaves it well behind cache-delta's reversible 43.8%. Reproduce: PYTHONPATH=. python benchmarks/codebench.py 20.

Historical: the 2026-06-23 measurement

Kept for context, not current — superseded by the 2026-09-04 run above (newer package versions, a larger sample, and the 1.49.0 provenance-exemption fix that changes what the bare digest row even measures). This table ran against headroom-ai 0.27.0 with 16 sessions / 256 turns.

methodtoken savings$ savings (cache-aware)latency / turnfidelity
distil (PAYG digest)91.5%91.1%0.08 msreversible
distil + cache-delta89.6%89.0%12.9 msreversible
LLMLingua-2 (llmlingua 0.2.2, real)56.8%57.2%274 mslossy
distil-verbatim + cache-delta34.9%43.8%13.8 msreversible
distil-verbatim (Tier-0 only)0.0%0.0%0.6 msreversible
Headroom (headroom-ai 0.27.0, real, default)22.4%−16.8%5.3 mslossy

Honest reading at the time: the Tier-1 reversible digest was the dominant lever (~91% cache-aware, reversible, 0.08 ms) — it already captured the re-reads, so cache-delta added little on top of it. LLMLingua-2 was genuinely strong on savings (57.2%) but lossy and ~3,400× slower. Headroom reduced tokens 22.4% yet cost 16.8% more, busting the prompt cache with its default read-lifecycle compression. This benchmark also caught two real distil bugs (a cache-monotonicity flip and a Tier-0 token inflation) — both fixed in v0.22.0; the numbers above are post-fix.


The distil benchmark corpus, warm — same 2026-09-04 re-run

The deterministic corpus gate (9 domains, no API key) run against the same package versions. A cold process shows Headroom skipping compression entirely — its Kompress model is still loading — which trivially "passes" the decision-equivalence check on requests it never touched; the table below is the warm re-run, after Kompress finished loading.

techniquetokens saved$ saveddecision-equivgate
distil-causal52.9%58.7%100%✔ PASS
Headroom (headroom-ai 0.37.0, real, warm)1.7%2.0%81%✘ FAIL

Cold-process artifact, for the record: the same run before Kompress finishes loading shows Headroom at a misleading 0.0% tokens / 100% "PASS" — it fails open while loading and passes through every request untouched, which trivially agrees with the ungated baseline. Preload before comparing. Raw output: benchmarks/results/2026-09-04/.


The offline standings — deterministic runner, zero API key

The companion to the live head-to-head above: the same gate run by the deterministic (structural) runner on a broader 64-trajectory corpus — fully reproducible by anyone, no key required. Ranked by cache-aware dollar savings (see methodology below). Certified means the method passed a statistical non-inferiority test and preserved 100% of decisions; anything less is disqualified, however much it saved. Want to test a specific tool? Register it through the reproducible --external seam and it's measured on the identical axes.

TechniqueFamilyTokens saved$ savedDecision equiv.VerdictFidelity
distil-causalcache-aware + causal pruning 80.5%81.5% 100%✔ certified · leaderlossy*
truncate-tail (structural baseline)sliding-window / truncation 78.7%79.6% 14%✘ fails gatelossy
distil-streamlossless + cross-turn dedup (fully reversible) 61.0%61.7% 100%✔ certifiedreversible
distil-losslesscache-aware lossless + structured fold + template mining 57.4%58.1% 100%✔ certifiedbyte-exact
summarize (structural baseline)abstractive / rolling summary 56.5%57.2% 39%✘ fails gatelossy
LLMLingua-2 (real pkg)llmlingua 0.2.2, per tool-result 54.9%54.8% 0%✘ fails gatelossy
Headroom (real pkg)headroom-ai 0.27.0, whole-conversation optimize=True 43.5%44.0% 61%✘ fails gatelossy
extractive-prune (structural baseline)extractive importance (LLMLingua technique family) 18.2%18.4% 77%✘ fails gatelossy
minify-alllossless minification 0.1%0.1% 100%✔ certifiedbyte-exact

*distil-causal drops context that ablation proves never changed a decision — not byte-reversible, but certified decision-equivalent. The three Distil operating points are all certified at 100%: distil-lossless (byte-exact: structured fold + template mining), distil-stream (adds cross-turn dedup of recurring tool output the cache can't reach — fully recoverable), and distil-causal (adds causal pruning). The lossy families are faithful reference implementations; bring your own tool with --external to add it to the table. Numbers: claude-opus-4-8 pricing, deterministic runner. Reproduce below. Unaudited: this run predates the claims gate and its raw output was not committed, so these figures are not tied to an artifact; treat them as indicative until re-run.

The corpus

To avoid any one tool's home-turf advantage, the comparison runs on 64 reproducible trajectories across 8 families deliberately spanning both regimes: structured/repetitive data (JSON record arrays, SQL rows, metrics, logs) where structural compaction pays off, and diagnostic/prose content (Kubernetes incidents, stack traces, RAG chunks, support transcripts) that conservative crushers protect and lossy methods mangle. Decisions are buried inside large tool outputs — as on real agents — so naive head/tail truncation drops them (14% equivalence). A large reference doc is re-read every turn (the recurring tool output prompt caching can't reach), which Distil's cross-turn dedup collapses and others re-bill in full.

Why this is the honest comparison

Same gate for everyone

No method is special-cased

Every technique — including Distil's own — is scored by the identical decision-equivalence + non-inferiority gate and the identical cache-aware cost model. The harness will happily rank a competitor above Distil if it earns it. It doesn't, because no other family combines lossless, cache-stability, causal pruning, and certification — but the door is open.

Faithful baselines

Best-form, not strawmen

The baselines are faithful reference implementations of the real technique families — sliding-window truncation, extractive importance pruning (the LLMLingua / Selective-Context lineage), abstractive summarization, naive minification — each in its best reasonable form. They genuinely remove tokens. They just can't prove they kept the decision.

Raw ≠ real

Two ways to "win" that don't count

A method can post a big raw token cut yet flip decisions — summarize cuts 56% but keeps only 39%; truncation cuts 79% but keeps 14%; extractive importance cuts 18% but keeps 77%. All disqualified. Or a method can shave tokens yet bust the prompt cache and cost more in real dollars. The benchmark prices both, so neither illusion survives.

Verify it yourself

Plug in a real tool

Don't trust our reference baselines? Register any installed compressor through the --external seam and it's measured on the identical axes. The claim "Distil leads on certified savings" is reproducible — and falsifiable. That's the point.

The equivalence dial

100% decision-equivalence is the default — but it's a setting, not a wall. Some teams will trade a little equivalence for deeper savings. Distil makes that trade explicit and bounded instead of hidden: you set a target, and the compressor spends a divergence budget — floor((1−target) × turns) turns — on the highest-value turns first, falling back to byte-exact everywhere else. You always know exactly what you traded.

$ distil frontier --corpus benchmarks/corpus_xl
savings-vs-equivalence dial  (runner=deterministic)

   target   achieved equiv  token savings   curve
----------------------------------------------------------------------
     100%             100%          58.1%   ████████████████
      95%             100%          58.1%   ████████████████
      90%             100%          58.1%   ████████████████
      80%              82%          62.9%   ██████████████████
----------------------------------------------------------------------
At 100% you get the certified-safe result. Relax the target and the budget is
spent on the highest-value turns — deeper savings, a known equivalence cost.

Honest notes: the dial's resolution is bounded by session length — on 5-turn trajectories the budget steps in 20% increments, so 95% and 90% round to the certified-safe point; longer sessions dial finer. The extra savings are real but modest (a principled risk knob, not a magic unlock), and the trade is always reported, never silent. Grade the frontier against the live model with --runner anthropic.

Reproduce it

# the bundled 8-domain standings, offline, zero API key
distil benchmark

# the 64-trajectory varied corpus above (from a repo clone)
python benchmarks/gen_corpus.py
distil benchmark --corpus benchmarks/corpus_xl --html standings.html

# verify against a REAL external compressor (list[str] -> list[str] over block texts)
distil benchmark --external mypkg.compressor:compress:MyTool

# grade with the live model instead of the deterministic runner
distil benchmark --runner anthropic --tokenizer anthropic
Honest scope. Numbers use the deterministic (structural) decision-equivalence runner — reproducible by anyone, no key required — on the reproducible 64-trajectory corpus generated by benchmarks/gen_corpus.py. ~80% is near the certified ceiling for this corpus: beyond it you start dropping decisions, and no honest tool can exceed that without lying — which is exactly what every disqualified row tried to do. For task-accuracy on a public benchmark (τ-bench, SWE-bench, GSM8K), ingest its traces into a corpus and run with --runner anthropic: the same comparison, graded by the live model. The harness is the deliverable; the corpus is swappable.

Live-model validation — on your own traffic

The live head-to-head above used our determined corpus. The same harness grades your traffic with no new code: every technique's compressed context judged by the real model, on real benchmark trajectories. The path is three commands:

# 1) record a public benchmark's agent trajectories as API requests (jsonl),
#    then convert them into a Distil corpus
distil ingest --input taubench_traces.jsonl --out corpus_taubench

# 2) run the SAME standings, but decisions graded by the live model
#    and tokens counted billing-grade
export ANTHROPIC_API_KEY=sk-ant-…
distil benchmark --corpus corpus_taubench --runner anthropic --tokenizer anthropic

# 3) the certified leaders now carry a live-model decision-equivalence verdict

What this buys: the equiv column becomes "did the live model make the same decision on the compressed context," graded on τ-bench / SWE-bench / GSM8K episodes instead of the synthetic corpus. The non-inferiority gate, the cost model, and the technique set are unchanged — only the judge and the data get more real.

We did this — see the live head-to-head at the top. On a realistic, decision-determined corpus (where byte-exact = 0% proves the model reproduces its own decision on identical context), Distil's causal-prune + lossless certifies 83.2% savings at a 0% live decision-change rate (2026-07-05, distil 1.10.1 vs llmlingua 0.2.2 and headroom-ai 0.27.0) — the aggressive operating point does hold when the decision is actually determined by the context, and the real competitor packages run through the very same live gate.
The honest boundary (why the corpus matters). On decision-ambiguous traffic the live model varies on identical context — on the original synthetic corpus its byte-exact change rate was ~50%, so nothing can certify and equivalence is undefined (it's the model's coin-flip, not the compressor's fault). The fix is not to spin the number — it's to certify on traffic where the decision is genuinely determined (as real agent steps are). That's why the determinism precondition is checked and reported, not assumed. Recalibrate on your own traffic; the guarantee is marginal over the calibration distribution, not a per-prompt promise.

The degradation curve — what each rung costs

Everything above answers a yes/no question: is the shipped strategy non-inferior? That is the right question for a gate and the wrong one for a decision, because it says nothing about the shape of the tradeoff on either side of the operating point. This is that shape, measured over the offline corpus with distil bench --curve. No API calls, no model, no cost — so it runs per-commit rather than once a quarter.

Fact recall plotted against token savings for each rung of the compression ladder. The none, tier-0 only, subscription and lossless rungs all sit at 1.0 recall and differ only in savings; the aggressive rung drops to 0.55 recall and is marked as not reversible.
RungSavingsRecallVisibleFacts lostReversible
none0.0%1.0001.0000yes
tier-0 only0.0%1.0001.0000yes
subscription0.0%1.0001.0000yes
lossless47.0%1.0000.7690yes
aggressive65.5%0.5510.551828no

Distil 1.51.1, 9-trajectory corpus, generated 2026-09-06. Recall is the macro average across domains, counting facts that are recoverable through an expand handle; visible counts only facts still readable in the forwarded text. subscription is the lossless-only path driven through the adapter's real verbatim branch, not a reconstruction of it. Every rung carries production's reject-if-bigger guard, so no rung can report savings the proxy would not take. Regenerate with distil bench --curve.

Three things the curve says that the gate cannot

The results JSON is written to benchmarks/results/curve.json and the chart to docs/assets/curve.svg, both stamped with the version and date they were generated from. tests/test_curve.py asserts the shape rather than the pixels: reversible rungs must hold full recall, and the lossy rung must save more while losing facts. If the ladder ever changes shape, that test fails before this page goes stale.


Reproduce in CI

Reproduce the Numbers

Every headline number has a command. This page shows exactly how each one is produced — what is synthetic, what is live, and what we ran but chose not to ship as the default (E7). None of it requires you to trust our word; it requires you to run the script.

Short version. Three gates run in CI on every push: distil bench (offline decision-equivalence), distil verify (byte-fidelity), and distil validate (adversarial real-path). Real-model grading uses benchmarks/prove.py. The head-to-head comparison is benchmarks/derc_live_compare.py. A cross-compressor invariant scorecard is python benchmarks/scorecard.py. Full methodology in docs/EVALUATION.md.

The three CI gates

Every push to main must pass all three gates before merging. They run in ~2 min combined, no API key.

Gate 1 — distil bench (offline decision-equivalence)

Runs the bundled 8-domain trajectory corpus through the decision-equivalence gate using the deterministic (structural) runner. The gate certifies that compression does not change the agent’s next action on any trajectory. This is the primary correctness gate.

# reproduce exactly — no key, no internet, ~5 s
distil bench

# or against the full 64-trajectory varied corpus (requires a clone)
python benchmarks/gen_corpus.py      # writes benchmarks/corpus_xl/ (seeded, deterministic)
distil bench --corpus benchmarks/corpus_xl

The deterministic runner is a structural stand-in: it checks whether the “decision-bearing” token sequence survives (not a live model call). This makes it reproducible by anyone in ~5 seconds. It is not a substitute for real-model grading — see Gate 3 and prove.py below for that.

Gate 2 — distil verify (byte-fidelity)

Checks that every Tier-0 and Tier-1 operation in the corpus is byte-reversible: compression followed by decompression returns the exact input bytes, and frozen history never mutates between turns. A failure here means a lossless claim is false.

# verify byte-fidelity across the bundled corpus
distil verify

# against a custom corpus
distil verify --corpus benchmarks/corpus_xl

Gate 3 — distil validate (adversarial real-path)

Drives the actual compression + proxy path against 12 hostile inputs (huge log blocks, unicode surrogates, deeply nested JSON, handle-injection strings, secret-looking content, malformed/None tool results, empty blocks) and asserts five load-bearing invariants on each:

InvariantWhat it checks
reversibilityEvery digest handle recovers its exact original bytes from the local RestoreStore.
reject-if-biggerA compressed block is never larger than its original (plus a 64-byte slack for stub metadata).
recency-exactThe most-recent tool result — the agent’s freshest output — is byte-identical after compression.
fail-openNo input, however hostile, makes the compressor raise an exception or the proxy return 5xx.
content-freeAfter a run, no prompt/response/tool text appears in any on-disk telemetry file (only hashes, sizes, counts).

12 cases × 5 invariants = 60 checks. All 60 must pass.

# run the adversarial gate
distil validate

Source: distil/harness.py. The gate is separate from distil verify (corpus byte-fidelity) and distil bench (non-inferiority): it exercises the code paths a corpus never hits — hostile inputs, streaming, marker injection.


Real-model grading — prove.py

benchmarks/prove.py removes the circularity of the deterministic runner by grading real agent traces with a real model. It runs four experiments (E1–E4) against actual trajectories:

ExperimentWhat it measures
E1 FrontierToken savings vs. decision-change rate per compression level.
E2 Certification coverageCertify at α on a calibration split, then measure the realized decision-change rate on a disjoint held-out split over many random splits. The certificate is sound iff empirical P(realized ≤ α) ≥ 1−δ.
E3 Distribution shiftLeave-one-domain-out: calibrate on all domains but one, test on the held-out domain (the exchangeability stress test).
E4 Downstream task successConverts per-turn equivalence into actual outcome: a trajectory keeps its result iff every decision is unchanged. Requires outcome labels (τ-bench reward / SWE-bench resolved).

Four grading backends are supported:

--runnerWhat it usesWhen
smokeOffline heuristic — non-evidentialPlumbing / CI, no key. Not evidence about real agents.
claude-cliThe claude -p CLI — your Claude Code subscriptionNo API key needed if you already use Claude Code.
openaiAny OpenAI-compatible endpoint (vLLM / Ollama / LM Studio)Free at scale with a local open model.
anthropicThe Anthropic API (ANTHROPIC_API_KEY)Billing-grade reference for published numbers.
# offline plumbing check only — no key, no download
python benchmarks/fixtures/make_fixtures.py
python benchmarks/prove.py --dataset fixtures --runner smoke --alpha 0.2

# real-model grading via your Claude Code subscription (no API key)
python benchmarks/prove.py --dataset tau --path runs.json \
    --runner claude-cli --model claude-haiku-4-5-20251001 --samples 3

# local open model via vLLM (zero per-call cost)
# vllm serve meta-llama/Llama-3.1-8B-Instruct --port 8000
python benchmarks/prove.py --dataset swe --path swe_trajs/ \
    --runner openai --base-url http://localhost:8000/v1 \
    --model meta-llama/Llama-3.1-8B-Instruct

The smoke runner is a non-evidential stand-in for plumbing checks. It is not evidence about real agents. See benchmarks/PROVE.md for the full runner table and what each backend actually proves.


The head-to-head corpus

The head-to-head runs on 64 trajectories across 8 families, generated deterministically with a fixed seed so the numbers reproduce exactly. The corpus spans both regimes deliberately so no tool gets a home-turf advantage:

Decisions are buried inside large tool outputs, as in real agents, so naive head/tail truncation drops them. A large reference document is re-read every turn (the recurring tool output prompt caching can’t reach), which distil’s cross-turn dedup collapses.

# generate the corpus (idempotent, seeded at 42)
python benchmarks/gen_corpus.py      # writes benchmarks/corpus_xl/

# run the head-to-head — distil + structural baselines, deterministic runner
distil bench --corpus benchmarks/corpus_xl

# add a real external compressor (list[str] -> list[str] over block texts)
PYTHONPATH=. distil bench --corpus benchmarks/corpus_xl \
  --external benchmarks.headroom_adapter:compress:Headroom

# live model grading
export ANTHROPIC_API_KEY=sk-ant-...
distil bench --corpus benchmarks/corpus_xl --runner anthropic --tokenizer anthropic

The full live head-to-head (distil vs. llmlingua 0.2.2 vs. headroom-ai 0.27.0, graded by claude-opus-4-8 majority-of-3) ran on 2026-07-05 and its raw output is committed at docs/paper/results/derc_live_compare.2026-07-05.log. Reproduce it:

pip install headroom-ai llmlingua
python benchmarks/derc_live_compare.py

The invariant scorecard

The same five invariants distil enforces on itself (distil validate) can be run against any compressor adapter. benchmarks/scorecard.py does this and produces a side-by-side table. It runs offline with no extra installs; headroom and llmlingua columns appear automatically if the packages are installed.

python benchmarks/scorecard.py      # prints table, writes benchmarks/scorecard.json

What the current run produces (distil + structural baselines; external adapters skipped when not installed):

                    distil  truncate@500  recency-w@500  keep-last-3  recomp-extr  sel-context
  ----------------  ------  ------------  -------------  -----------  -----------  -----------
  reversibility     PASS    FAIL          FAIL           PASS         FAIL         FAIL
  reject-if-bigger  PASS    PASS          PASS           PASS         PASS         PASS
  recency-exact     PASS    PASS          PASS           PASS         PASS         FAIL
  fail-open         PASS    PASS          PASS           PASS         PASS         PASS
  content-free      PASS    n/a           n/a            n/a          n/a          n/a

Fairness semantics used per cell:


The negative result: E7

E7 is an internal SWE-bench Verified end-to-end experiment (benchmarks/swe_bench_e2e/) that ran the full agent loop — not a single-step proxy — under distil’s aggressive lossy compression tier. The result: 52% → 16% task resolution. The reversible tier held (56% vs. 52% full-context baseline). Per-step decision-equivalence had passed; end-to-end task success still cratered under the aggressive tier.

We publish E7 as evidence, not confession. It is the concrete reason the shipped default is the conservative/reversible tier rather than the tier that scores best on a raw compression-ratio leaderboard. The trajectory-risk certificate and the fail-safe gate exist because of E7.

# E7 harness (requires SWE-bench Verified setup)
bash benchmarks/swe_bench_e2e/run_all_agents.sh
bash benchmarks/swe_bench_e2e/run_all_scores.sh
python benchmarks/swe_bench_e2e/aggregate.py

Full methodology, the A/A nondeterminism baseline, trajectory-risk certificates, and what distil’s numbers do and don’t prove: docs/EVALUATION.md.


Synthetic vs. live: what each number actually is

NumberHow producedWhat it provesWhat it doesn’t
83.2% savings / 0% decision-change (synthetic 120-turn corpus, live-graded head-to-head, 2026-07-05, distil 1.10.1 vs llmlingua 0.2.2 and headroom-ai 0.27.0) Real packages, real model (claude-opus-4-8 majority-of-3), 120-turn corpus On this corpus, distil preserves agent decisions at that compression level End-to-end task success on a different workload; E7 is the cautionary data point
80.5% certified savings (offline standings) Deterministic structural runner, 64-trajectory varied corpus Decision-bearing token sequences survive at this compression level, reproducible by anyone Live model agreement; uses a proxy for the model’s decision, not the real model
60/60 validate checks Real compression path, 12 adversarial inputs, 5 invariants The five guarantee properties hold on hostile inputs distil’s unit suite never exercises Compression ratio; orthogonal to the savings numbers
SWE-bench Lite: −2.0 / −1.7 pts, cost-neutral (shipped path, n=300, two runs) Plain vs distil-served Claude agent on all 300 tasks, official grader, prompt caching on both arms The shipped path is within about 2 points on task success and does not cost more on these tasks Non-inferiority at the 5-point margin (not shown); long sessions (not measured); research configurations (E8–E14) say nothing about the shipped path
E7: 52% → 16% (aggressive tier) Same SWE-bench Verified setup, aggressive lossy tier Aggressive compression with a passing per-step certificate still cratered end-to-end Anything about the reversible / gated tier: that one held at 56%

Independent validation

What an outside benchmark found

In August 2026 an engineer ran distil against bare Claude Code across 75 agent sessions and measured the result from the API’s own usage fields rather than from any tool’s dashboard. One distil configuration finished 6 of 15 coding tasks where bare Claude Code finished 13. This page is what they found, why it happened, and what changed in 1.49.0.

We publish it for the same reason we publish the bugs our own evaluation harness had: a compressor whose public record contains no failures is not a compressor that never failed, it is one that never looked. The measurement that matters is the one taken by someone with no stake in the answer.

The result

ArmTasks completedCache writes vs baseline
bare Claude Code (baseline)13 / 151.00×
distil hook mode12 / 150.94×
distil proxy, default (subscription)12 / 151.56×
distil wrap --expand6 / 150.87×

Cache writes are the number that decides cost on a cached workload, and they are taken from cache_creation in the provider’s response — not from distil’s own accounting, which is the thing under test.

The failure that mattered

Seven --expand runs ended with an empty diff: the agent ran for 18, 22, even 51 turns, wrote nothing to disk, and reported success. On the long refactor it did this all three times — and the repository’s 1,371 existing tests still passed, because nothing had been modified. A benchmark that graded on “did the suite stay green” would have scored those runs as wins.

No other arm produced a single empty diff in 60 runs.

Cause 1 — the agent never ran its own tool call

Claude Code emits parallel tool calls, so one assistant message routinely carries both an Edit and a distil_expand. distil’s recovery loop assumed distil_expand was the only tool call in the turn. It answered just that block and re-queried — which left the client’s Edit as an unanswered tool_use upstream, and let the continuation’s stop_reason: end_turn replace the turn’s real tool_use.

Claude Code executes tools only when stop_reason == "tool_use". So it received the Edit, ran nothing, printed the continuation’s “all done”, and finished. Fixed: a turn carrying a client tool call is now terminal — relayed verbatim, stop_reason intact, no re-query.

Cause 2 — a failed recovery truncated the stream

When the re-query failed, the connection closed with no message_delta and no message_stop. A terminator-less SSE message is truncated, not finished, so the SDK retries it invisibly: wall-clock burned, no progress, and no error the user can see. That is the benchmark’s 808-second, zero-write signature. Fixed: every exit emits a terminator.

Cause 3 — a file read could be digested before it was edited

Recency is positional, so a file read three turns ago was eligible for digesting. But the agent must still reproduce that text character-for-character in an Edit(old_string=…). Once the read is a digest, no exact match exists. Fixed: the exemption is now keyed on provenance — results answering Read/Grep/Glob and their MCP equivalents stay byte-exact at any age. Logs and test output still compress normally.

Compression that saved nothing still cost cache

On a subscription distil stays lossless-only by design and correctly reported 0.0% savings — while still writing 1.56× the baseline’s cache-creation tokens, and 2.52× on a short session. Doing nothing to the content is not the same as doing nothing to the cache.

Two causes, both fixed. The distil_expand tool was injected only once a recoverable handle existed, so the tools array — which Anthropic caches ahead of the system prompt and all history — changed shape on the turn compression first fired. And an unmodified request body was re-serialized rather than forwarded byte-for-byte; json.dumps is not a byte-faithful round-trip, so the prefix moved even when nothing had been compressed.

Worse, distil’s own drift report sent users hunting in the wrong place: it blamed “a tool list whose order varies” upstream, while distil’s tool list was the one varying. It now names distil’s own causes first.

The mode that won had nothing to show for it

Hook mode was the strongest distil arm measured — 12/15 tasks, 0.94× baseline cache writes, the highest cache-hit rate of any configuration — and the only one that wrote no ledger at all. Its effect was visible solely in the provider’s billing, which is an awkward position for a product whose claim is that it proves its numbers instead of asking to be trusted.

It now writes a content-free receipt per compressed result:

distil hook --install     # documented Claude Code extension point
distil hook --stats       # what it actually saved

What this does not establish

The benchmark’s own caveat is worth repeating: an empty diff proves the agent stopped early, not that compression caused it. The runs used --shadow 0.0 to conserve tokens, so distil’s decision-equivalence sampler — the tool built for exactly this question — captured nothing. We have since reproduced the mechanism locally, without API spend, and pinned it with regression tests that fail without their fixes. But the association came first and the proof came second, which is the right order to admit to.

The lesson we are keeping: a green test suite does not prove the work was done. Acceptance has to check that the required change exists, not merely that nothing broke.

Related: the bugs our own harness had · negative results · why removing tokens can cost more · what “savings” means on a flat-rate plan


Evaluation

Every compressor asks you to trust it. This page is what distil measures instead — including the numbers that make distil look worse, because a suite that only ever produces good news is a suite nobody should believe.

Why compression ratio is not a result

Token reduction and task success are not linearly related, and the relationship can cliff-edge. One study found 4× compression collapsing SWE-bench resolution from 86% to 7% — not a decline, a cliff.

We hit it ourselves. E7 ran the full SWE-bench Verified agent loop under distil's aggressive tier and measured real resolution: 52% → 16%. The reversible tier held (56% vs 52% full-context). Per-step decision-equivalence had already passed. End-to-end task success still cratered.

We publish E7 as evidence, not confession. It is why the shipped default is the conservative tier rather than the one that wins a compression-ratio leaderboard.

The rule that follows: every compression claim reports token savings and a task-success delta together, at the setting we actually ship. A ratio with no success number beside it is an invitation to find the cliff in production.

Distil's evaluation stack: five per-commit gates — bench for decision equivalence, verify for byte fidelity, retention for fact recall, fidelity for state fidelity, validate for adversarial invariants — under a conformal risk certificate, with external-validity checks and a versioned record for every run.

Five layers, five different questions

LayerQuestionCommandCost
Decision equivalenceDoes the agent's next action change?distil benchfree
Byte fidelityIs every compression exactly reversible?distil verifyfree
Fact recallWhich facts stay visible, recoverable, lost?distil retentionfree
State fidelityDoes a yes on recall actually mean anything?distil fidelityfree
AdversarialDo the invariants hold on hostile input?distil validatefree

All five are offline, need no API key, and run in seconds. That is deliberate: a gate you skip because it costs money is not a gate. All five run per-commit in CI.

The case recall cannot see

A file created at turn 2 and deleted at turn 4. When compression drops turn 4 every path token remains, so string recall reads 100 percent, but artifact-state fidelity reads 0 percent — the failure is stale (present but wrong) rather than lost.

Recall asks "is the fact still there". Here is a trajectory where the answer is yes and the agent is still broken:

turn 2   Write(file_path="net/scratch_bench.py")
turn 4   rm net/scratch_bench.py

Compress away turn 4. Every path token is still present — string recall reads 100% — while the agent now believes a file exists that does not, and will plan around it. Nothing in the transcript says otherwise.

distil fidelity folds tool calls into a file-state ledger and grades the final state, splitting two outcomes that presence metrics conflate:

OutcomeMeaningSeverity
exactfinal state preserved—
lostpath absentloud — the agent can see the gap
stalepath present, wrong statesilent — the agent acts on a false belief

These are never averaged into one accuracy figure, because a compressor that drops a whole file history is safer than one that preserves half of it. On the case above: string recall 100%, state fidelity 0%.

Factory.ai measured every method they tested at 2.19–2.45 / 5.0 on this axis across 36,611 production engineering messages. Presence-based metrics are structurally unable to detect it — the string is present.

Three more things recall misses

Overclaim. "approximately 4200 ms" → "4200 ms". Byte-identical value, and every recall metric scores it perfect — but the agent has been handed a precision the source never asserted. The hedge was the information. Hedges are grouped into classes so reshaping (approximately → about) is not penalised; only the disappearance of hedging is. Direction is asymmetric: overclaim is gated, underclaim only reported.

Continuation. Whether the agent still knows what is left to do. Dropping a completed item is cheap — work gets redone. Dropping a pending one is silent: the work is skipped and success is reported anyway.

Propagation. Whether a loss at turn k shows up as a behaviour change at turn k+n, as a lag-lift profile. Two limits are printed in the tool's own output: it is association rather than causation, and periodic workloads alias at multiples of their period.

The numbers, on the shipped tier

$ distil fidelity
artifact-state fidelity   100.0%  (7/7 artifacts intact)
  stale (wrong state)       0.0%  (0)
  lost  (absent)            0.0%  (0)

hedge fidelity             94.7%  (162/171 claims kept their hedging)
  overclaimed               5.7%  (9)  <- value kept, uncertainty dropped

pending-work recall       100.0%  (2/2 remaining items still visible)

error propagation  (36 turns, base decision-change rate 0.0%)
  verdict: no decisions changed at all — nothing to propagate, and nothing tested

OUTPUT surface — past answers digested on re-entry (the other half of the bill)
  (6 blocks changed by digestion, 4 facts removed from them — the graded evidence)
  artifact state         100.0%  (7/7)   stale 0
  hedge fidelity         100.0%  (6/6)   overclaimed 0
  pending-work recall    100.0%  (2/2)   dropped 0

silent failures, BOTH surfaces: 9  (input 9, output 0)

Those 9 overclaims are real, and we ship the number rather than tune it away. Tier-1 digests spans behind restore handles, and a digested hedged sentence can keep its value while losing its qualifier. Unlike a lost fact this is not recoverable in practice: a missing fact prompts the agent to expand, a missing hedge gives it no reason to look.

So the CI gate is --max-silent 15, not 0. Gating at zero would assert a property the compressor does not have. The band fails a regression that doubles it — and fails an improvement too, forcing the number to be re-read rather than drifting.

Public benchmarks — falsifiable by a stranger

Every other number on this page is graded against distil's own corpus. Rigorous, and unfalsifiable by someone who doesn't trust us. distil suite closes that: twelve public benchmarks whose answer keys were written by other people.

It costs nothing to run. Grading is deterministic recall against the answer key — no model in the loop, no API key, no spend. Suites that grade with an LLM judge cost real money per tier, which makes them something you run before a launch rather than before a merge. This one is wired into make gate and the CI gate job.

Two benchmark cases compared: a thin-payload GSM8K question with nothing to compress, labelled a control, beside a rich-payload BFCL tool schema whose compression result is evidence.

The payload column is the important one. A GSM8K case is a one-line word problem: there is nothing in it to compress, so an unchanged score proves the compressor left it alone. That is a control. Only rich-payload rows — tool schemas, retrieved passages, long-form narrative, code — can demonstrate compression quality. A run that grades only controls exits 1.

bfcl leads tier 1 deliberately. Berkeley Function Calling compresses the tool schema and checks that every name the gold call is built from — the function and each argument it passes — survives compression: the failure an agent proxy is most likely to cause and least likely to notice, because no QA benchmark ever asks the model to act. The gold call itself is not graded as an answer — it never appears in the schema text, and checking whether a model still emits it would need a model in the loop, which would stop this suite running in the CI gate. What is measured is stated exactly.

uv run distil suite --tier 1              # tool-calling + retrieval
uv run distil suite --tier 1 --tier 2     # add the harder payloads
uv run distil suite --only bfcl -n 50     # one benchmark
uv run distil suite --json --max-lost 0   # gate

A benchmark that cannot be fetched is reported as a FAILED row and exits 1 — never dropped. A suite that silently skips what it could not load reports a clean sheet for a run that measured less than it claimed.

Every run is a record, not a number

--json emits an envelope, validated against a published JSON Schema on every test run:

{
  "schema":  "distil.eval/1",
  "subject": { "compressor": "serving", "module": "distil.compress.strategies" },
  "dataset": { "trajectories": 9, "fingerprint": "sha256:051b836358932883" },
  "grader":  { "kind": "deterministic",
               "detail": "synthetic DECISION: oracle — NOT a model" },
  "metrics": { ... },
  "gates":   [ { "name": "max_silent", "threshold": 15, "observed": 9,
                 "passed": true, "rationale": "..." } ],
  "passed":  true
}

On your own traffic

The offline gates are graded on our corpus against our oracle — rigorous, but not checkable by you. Two things close that:

# graded against a PUBLIC benchmark's answer key, not ours
distil retention --dataset hotpotqa

# A/B on YOUR traffic, with an A/A control
distil wrap --shadow claude
distil shadow-stats --record

Shadow mode replays the same compressed request to measure the grader's own disagreement with itself (A/A), then subtracts that floor from the A/B divergence. Without it you would read sampling noise as compression harm. --record emits the same envelope as the offline gates, so a live result is as attributable as a corpus one — with a traffic-window descriptor in place of a content hash, because live traffic has none and must never grow one.

Everything here is content-free. The meters store counts; no prompt, path or tool output leaves your machine.

What the probes found in our own work

An eval suite's real test is whether it catches its authors. These are recorded in full in the paper:

Each of those is a failure the probes exist to detect, occurring in the probes. We publish them because the alternative — a suite that has never been wrong — is the less trustworthy artifact.

Who grades the live certifier

A live certificate is graded by a model, so changing that model changes what the certificate means. Since 1.56.0 the default live certifier is claude-sonnet-5-5 at effort=low, chosen by a pre-registered, paired migration eval on real τ-bench traffic rather than by price. It passed every gate, including a quality lower bound of at least −5 points against claude-opus-4-8, at 57.5% lower cost per case; claude-opus-5-5 did not clear the quality bound and claude-haiku-4-5 was too weak. --effort is now a flag on certify, eval, benchmark, frontier and conformal. The earlier published live result was graded by claude-opus-4-8 and stays labelled so.

The same harness exposed a gap in what the certificate covered. The certified distil strategy digests only the latest tool output and saves 2.6% of tokens on real SWE-agent trajectories; the serving adapter digests every earlier tool output and saves 52.6%, so most served bytes were never certified. distil certify --strategy served now runs the real adapter. Its result is an open risk: with recovery it keeps the next action in 78.0% of held-out coding decisions, against 94.1% for two identical uncompressed calls. Whether that costs solved tasks is the question of a SWE-bench outcome eval that is built and tested offline and has not been run.

Bubble chart of each certifier setting on held-out tau-bench turns: cost per case across, expand action-equivalence up, bubble area for self-consistency. The cheaper claude-sonnet-5-5 at low effort sits highest, and claude-haiku-4-5 is cheapest but lowest.

Bubble chart of the SWE-bench Lite outcome runs, one bubble per run and arm: cost per task across, share of tasks resolved up, bubble area for the number of paired tasks, coloured by plain or distil-served agent.

Two harness lessons belong with the probe list above: a distil-wrapped shell exports ANTHROPIC_BASE_URL to the local proxy, so an eval client must pin api.anthropic.com; and the bundled corpus plants DECISION: markers for the offline oracle, so live grading needs marker-free real traces. The method, tables, diagrams and rerun commands are on Model Migration.

What the exit summary proves

Every distil wrap session ends with a proof ledger, and the same four verdicts appear in distil stats and distil dissect — one function behind all three, so two surfaces reading one ledger cannot disagree about it. Each line is allowed to come back negative. That is the only reason any of them is worth printing.

  distil proof ledger — session 47m        # the four lines, with YOUR traffic's numbers
    budget   decision-change budget: intact (e-value <E>, n=<N>) — bound within the <B>% budget
    risk     decision-change risk ≤ <P>% (95% conformal bound, n=<N>) — within the <B>% budget
    output   the model's replies were <T> tokens shorter per request under compression
             (95% CI [<lo>, <hi>], n=<N>, shadow-measured on all traffic; not --shape-output)
    receipts <N> receipts, chain intact — <k> re-checked since this machine's last pass; `distil receipts` re-hashes all

No sample values are shown here on purpose: these are the only numbers on this site that are not published measurements, because they are yours. Each line can also read BREACHED at sample k … compression held at lossless-only, unproven … the bound is above the budget, ABOVE the budget, chain BROKEN at receipt k, or no measurable effect on reply length.

Below the reporting floor, every line withholds its number and prints how far along it is instead — not enough samples yet (12/50). A verdict computed over evidence too thin to support it is worse than no verdict, because it teaches the reader to ignore the line.

The numbers above are the shape of the output, not a claim. The maintainer's own sample — 398 A/B replays against 399 A/A controls, published content-free — puts compression at 2.5 pp of decision agreement, 95% CI [0.5, 4.5] pp, inside the 5% budget, with replies 68.3 tokens shorter per request, 95% CI [-147.2, -2.9]. Your traffic is not that traffic, which is the entire reason the lines are computed on yours.

Run it yourself

$ uvx --from distil-llm distil bench      # ~10s, no API key
$ uvx --from distil-llm distil fidelity   # the state probes
$ make gate                              # everything CI runs

Full methodology in EVALUATION.md; step-by-step instructions in RUNNING-EVALS.md.


The Corpus

The asset that makes certification meaningful — 9 real, captured-style agent trajectories across 9 domains. Every strategy must pass the gate on every one of them.

Bar chart of lossless savings across 9 domains, all passing the decision-equivalence gate: ops/sre 32.8%, coding 25.5%, support 32.6%, research 25.7%, data-analysis 18.1%, devops 22.8%, finance 24.9%, web-research 89.8%, agent-worklog 35.3%; aggregate 48.4% cheaper reversibly across 7,080 prunable tokens; figures repeated in the table below

Why a corpus matters

A compression strategy is not trustworthy because it works on one example. It needs to hold across the full breadth of what agents actually do: incident response, code debugging, customer support, research synthesis, data analysis, infrastructure work, and financial reconciliation.

The rule: a strategy that fails the TOST non-inferiority gate on any trajectory does not ship. distil bench runs this as a CI gate and exits non-zero on failure.

Domain table

Domain Trajectory ID $ Saved Distil Aggressive Prunable tokens
ops/sre sre-disk-incident 32.8% PASS FAIL 615
coding coding-bugfix 25.5% PASS FAIL 736
support support-refund 32.6% PASS FAIL 765
research research-synthesis 25.7% PASS FAIL 809
data-analysis data-analysis-sql 18.1% PASS FAIL 965
devops devops-rollback 22.8% PASS FAIL 857
finance finance-reconcile 24.9% PASS FAIL 1014
web-research web-research 89.8% PASS FAIL 428
agent-worklog agent-worklog 35.3% PASS FAIL 891
Aggregate 48.4% PASS FAIL 7,080

Numbers are reproducible from the bundled corpus using the heuristic tokenizer. Absolute dollar figures assume claude-opus-4-8 public list pricing — verify before billing use.


Trajectory structure

The bundled trajectories are 4-turn headless agent loops (the structural minimum, enforced by corpus.validate(), is ≥3 turns). Three structural components:

Trajectories are stored as JSON files in corpus/. The manifest is at corpus/manifest.json. The loader (distil/corpus.py) resolves the corpus path from the wheel, the repo, or a $DISTIL_CORPUS environment variable.


distil bench — the CI gate

distil bench is the CI gate for new compression strategies: it runs the full corpus measurement loop and exits non-zero if any of the following occur:

$ distil bench
# Exits 0 only if every trajectory passes all three conditions above

GATE: PASS — every trajectory certified non-inferior; aggressive rejected on all.

# If a new strategy were added that degraded quality on one domain:

GATE: FAIL (1 issue(s))
  - corpus/finance-reconcile.json: distil FAILED non-inferiority

The rule in one sentence

A new compression strategy must pass make gate (non-inferior on every domain, byte-reversible). No green gate, no merge.

Corpus invariants enforced by corpus.validate()

The validate() function checks structural correctness of each trajectory before it enters the gate:

A trajectory that fails validation is reported as a failure in bench before any cost or quality measurement runs — there is no silent skip.