compression with a quality contract

Session-level A/B

Every other savings number distil prints is per request: tokens removed from what was sent. That cannot tell you what a session cost, and it cannot survive a model change. distil ab can. It compares sessions distil compressed with sessions it deliberately left alone, chosen at random, running at the same time, on the same model. Individual results take months to become conclusive, and the report says how many sessions it still needs.

Why per-request savings are not enough

A new model version changes how an agent works. It may take more turns to finish the same work, or fewer. All of that moves your bill, and none of it is distil. If you compare last month with this month, the model change lands in the comparison as if distil had caused it. In the simulation behind this page, a model that needs about 40% more turns makes a before/after comparison of the distil sessions read 40% more cost, while distil was saving 20% the whole time.

The only comparison a model change cannot contaminate is one where both sides see the change at once. So distil holds out a small random share of sessions and compares each group only with the other, at the same time, on the same model.

The holdout

What is measured

The headline is mean cost per randomised session, distil against holdout: the provider's own usage priced at list, cache reads at a tenth and cache writes at 1.25 times the input rate. Every call distil made on the session's behalf counts, including a distil_expand re-query and a shadow replay. Every randomised session that has ended counts, including one that failed (at what it actually cost, usually nothing). A session on a model distil cannot price is counted and left out.

Beside it are the mediators: turns per session, tasks per session, cost per task and cost per turn. A task is the stretch between two things you typed. These are outcomes distil can change. If compression made you re-ask more often, cost per task would fall even while each session cost more, so a mediator is never the headline. In the simulation of exactly that case, distil making sessions 10% dearer with 25% more re-asks, the headline read +9.2% and cost per task read about −12%.

A balance check prints how each group's sessions ended: normally, in an error, or abandoned. A failure mode distil causes shows up there first.

Staying within one model

Sessions are grouped by the model their agent loop runs on, which is known at the first request and cannot be affected by compression, and by the client and its version (for example claude-cli/2.1). When a newer version of the same model family appears, the headline compares within the new model only, and says so: model changed on 2026-08-05: comparing within the new model only. The old model's result is kept, and the two are compared as a difference-in-differences.

A model can also change behind an unchanged id. Distil watches the holdout group's cost per session for a sustained shift, and when one is detected it starts a new era. It watches the holdout group because that is the one series distil cannot move, so a distil upgrade never looks like a model change.

The part that actually protects the estimate is the randomisation, not the change detection. Both groups run at the same time, so a model change moves both and cancels in the contrast.

Reading the report

distil ab — does distil lower what a session costs? (randomised holdout)
  A/B holdout: 5% of new sessions run uncompressed so distil can measure what it saves per session, causally (distil ab). Opt out: distil ab --holdout-rate 0, or DISTIL_HOLDOUT_RATE=0.
  Scope: this machine's own sessions only.
  window: all time (the anytime guarantee holds within the current era) · 7500 sessions (ended+priced: 7122 distil, 378 holdout) · 0 open · 0 unpriced

  mean cost per session  -11.8%  (95% anytime CI -36.2% … +33.0%; not significant)
    the one primary claim: empirical-Bernstein CS, cost capped at $500/session; anytime-valid within this era, no normality assumed
    Individual results take months to become conclusive: need ≈43,360 sessions (2,168 held out).
  bootstrap cross-check (fixed-n, strata-pooled): -21.2% … +17.6%
  efficient estimate (strata-pooled, CUPED, asymptotic)  -9.9%  (95% CI -34.3% … +23.4%; e-value 0.69; not significant)
    CUPED variance −5%
  mediators — outcomes distil can change, so never the headline:
    turns per session  +0.4%  (95% CI -5.6% … +6.7%; e-value 0.0949; not significant)
    tasks per session  —
    cost per task      -6.2%  (95% CI -32.1% … +29.6%; e-value 0.537; not significant)
    cost per turn      -6.6%  (95% CI -32.3% … +28.8%; e-value 0.547; not significant)
    (secondary claims are asymptotic and not multiplicity-adjusted; only the headline spends the error budget)
  cost of the holdout: those 227 sessions cost $38,668.18; compressed, ≈$4,550.20 more (range −$12,770.40 … $13,983.17)
  model changed on 2026-08-05 (claude-opus | claude-cli): comparing within the new model only

  balance check — how sessions ended, per arm:
    distil   capped 452, ok 7122
    holdout  capped 34, ok 378
    capped = sessions above the $500 cap, counted at the cap

  strata (model | client, era)                  n dist  n hold  $/sess dist  $/sess hold  Δ $/session  Δ turns/sess  Δ $/task
  claude-opus-4-8 | claude-cli/2.1 #1             2849     151      $147.05      $171.81       -10.9%         +0.4%    -14.4%
  claude-opus-5 | claude-cli/2.1 #1*              2854     146      $166.95      $184.01       -11.8%         +0.6%     -9.3%
  claude-sonnet-4-6 | claude-cli/2.1 #1*          1419      81      $145.80      $145.71        -6.2%         +0.1%     +0.1%
  * = current era, pooled into the headline · Δ columns are per-cell, CUPED-adjusted and asymptotic

  eras / change points
    2026-08-05  claude-opus | claude-cli: model changed claude-opus-4-8 → claude-opus-5

  difference-in-differences across a rollout ((distil−holdout)_after − _before; asymptotic)
    claude-opus | claude-cli: claude-opus-4-8 | claude-cli/2.1 → claude-opus-5 | claude-cli/2.1: -1.0%  (95% CI -61.6% … +155.2%; e-value 0.766; not significant)

That is a simulated history at the default holdout rate, with session costs about as spread out as real ones. Opus changes version on 5 August and the new version needs 40% more turns in both groups. After 7,500 sessions the headline is still not distinguishable from zero, and the report says how many sessions it needs. That is the honest answer at a small holdout. The difference-in-differences across the rollout is near zero, because distil's effect did not change.

Cost of the holdout is what the held-out sessions would have saved had they been compressed. It is the price of knowing. A larger holdout gets an answer sooner and costs more.

The math

The headline interval

Session costs are heavy-tailed: on the maintainer's own machine their spread on a log scale is about 1.9. An interval that assumes a bell curve fails there. So the headline makes no such assumption. Each group's per-session cost is capped at a limit fixed in advance, $500, and the report counts the sessions it touched. An empirical-Bernstein confidence sequence (Waudby-Smith and Ramdas, 2023) is built for each group's mean, each at half the error budget. The two combine into an interval on the ratio. Bounded data is all it assumes, and it is valid however often you look: the chance it ever excludes the truth is at most α, distil's single failure budget (distil.conformal.BUDGET_DELTA).

Capping has a price: the estimand becomes the capped mean. The capped group with the dearer sessions loses more, so a real saving reads smaller, never larger.

HoldoutSpread (log-sd)Headline: false alarm, 20 looksPrevious interval (mixture SPRT): false alarm
5%0.60.0%4.0%
5%10.0%6.5%
5%1.50.0%10.5%
5%20.0%10.5%
50%0.60.0%2.0%
50%10.0%2.5%
50%1.50.0%1.0%
50%20.0%0.5%

Each row is 200 simulated histories of 4,000 sessions with no real effect, checked 20 times. The previous interval, a normal-mixture sequential test, exceeded the budget once the tails got heavy at the shipped rate. The headline did not. The cost is power: a true 20% saving at the shipped rate and log-sd 1.0 was detected 1.5% of the time after 4,000 sessions and 76% after 20,000.

Secondary estimates

Beside the headline the report prints an efficient estimate of the same quantity. It pools the model-and-client cells by inverse variance, reduces variance with CUPED (each session's cost regressed on its project's cost over earlier sessions, which were assigned independently), and uses the normal-mixture sequence. It is sharper and only asymptotically valid: in the simulation, where projects differ a lot, CUPED cut its variance by 44%. The mediators, the per-cell table and the difference-in-differences use the same machinery. None of them spends the error budget, and none is adjusted for multiple comparisons. A percentile bootstrap is printed as a fixed-sample cross-check. When it and the headline disagree on whether the effect is distinguishable from zero, the report says so.

The anytime guarantee holds within the current era. --window slides a fixed window over time, which makes each view a fixed-sample look, and the report labels it that way.

Detecting a silent change

A two-sided self-starting CUSUM runs on the holdout group's log cost per session (k = 0.5, h = 12). A new era starts at the first session after the alarm, and the sessions between the estimated change and the alarm are dropped. It reads the holdout group on principle: distil cannot move that series, so an upgrade to distil is never mistaken for a model change. At the shipped rate the choice barely matters. Under a silent model change, reading both groups, reading the holdout, and running no detector at all land within noise of each other. An earlier version read both groups and claimed a measured gain from it, which the tracked simulation did not reproduce. The threshold is high because at a small holdout every false era throws away scarce data. With a model change hidden behind an unchanged id, the headline read +1.5 points above the true saving (±0.5); with an explicit version change, +1.9 (±0.5). Both are the cap at work: sessions averaged $150 to $210 against the $500 limit, and capping pulls a saving toward zero.

What it does not do

Every measured number on this page is reproduced offline by benchmarks/abtest_montecarlo.py into benchmarks/results/abtest-montecarlo.json. The decision record is ADR 0019.