compression with a quality contract

Model migration

A distil certificate is graded by a model. Change that model and you have changed what the certificate means, so distil does not change it on a hunch. This page is the eval that decided the default live certifier, the gates it had to clear, what it found about the served path while doing it, and how to rerun all of it. One result is good news and one is an open risk; both are here.

The question

Moving the grader from an incumbent model to a newer one is a migration, and a migration has one test: given the same contexts, compressed and not, does the new model reach the same verdicts? The verdict distil cares about is narrow. Is the agent's next action unchanged by compression, when the model is allowed to recover a digest with distil_expand? That is the metric distil is deployed on, and the one this eval treats as the headline.

The eval asks it twice, about two different things. First, which model should grade (the certifier result). Second, does the path distil actually serves to a caching client stay inside the same contract (the served result). The second question only arose because the first eval's harness would not let it stay hidden.

Method

Each case is one decision point from a real trajectory: a τ-bench turn, or a SWE-agent step. Per case and repetition, distil's own runners answer the same question under six arms.

Data. Public τ-bench historical trajectories (airline and retail, GPT-4o and Claude Sonnet agents), SWE-agent GPT-4o trajectories from the SWE-bench Lite leaderboard submissions, and SWE-bench Lite itself for the task-outcome runs. Every download is pinned and hash-verified; the full table, with files, selection seeds and pins, is in EVALUATION.md §6.9.

One trajectory decision point replayed on six arms: full context twice as a noise floor, the certified distil strategy, the same with distil_expand recovery, the served adapter, the same with recovery, and a truncation control that should diverge. All six go to one certifier, and the eval reports paired metrics: equivalence, self-consistency, control sensitivity, paired delta with interval, and cost per case.

Cases are paired: each candidate grades exactly the cases the incumbent did, each case averaged over its repetitions, and the difference carries a 95% interval from the case-level differences. Cost is the provider's own usage at list price over every call. Candidates are compared on held-out cases, set aside before any candidate was run. A harness hash gates every run, so a change to the runners or prompts has to be approved before its numbers can be mixed with older ones.

Two lessons came out of building it, and both would have ruined a result silently:

Which model certifies: the result

The real-traffic set is 100 τ-bench decision points, 60 held out, each run 4 times. The incumbent was claude-opus-4-8. Candidates were the newer claude-opus-5-5 (default effort, and low), claude-sonnet-5-5 at low effort, and claude-haiku-4-5.

CertifierNext action kept, with recoveryPaired change, lower bound (pts)Distil alone, no recoveryTwo full calls agreeControl divergesCost per case
claude-opus-4-8 (incumbent)90.8%reference61.3%92.5%89.2%$0.1458
claude-opus-5-590.8%+0.0, −10.073.3%93.3%73.3%$0.1336
claude-opus-5-5, effort low90.0%−0.8, −9.875.0%94.2%75.8%$0.1256
claude-sonnet-5-5, effort low94.6%+3.8, −4.566.7%99.2%79.6%$0.0619
claude-haiku-4-585.0%−5.8, −15.844.2%96.7%93.3%$0.0248

Paired change in the deployed metric against claude-opus-4-8 with 95 percent intervals, and cost per case. Only claude-sonnet-5-5 at low effort has an interval whose lower bound clears the gate of minus 5 points; claude-opus-5-5 at two effort settings and claude-haiku-4-5 do not. claude-sonnet-5-5 at low effort costs 57.5 percent less per case than claude-opus-4-8.

Bubble chart of each certifier setting on held-out tau-bench turns: cost per case across, expand action-equivalence up, bubble area for self-consistency. The cheaper claude-sonnet-5-5 at low effort sits highest, and claude-haiku-4-5 is cheapest but lowest.

The default is now claude-sonnet-5-5 at effort=low. It was accepted under five gates written down before the run, and it passed all of them:

Read it for what it is. The point estimate is above the incumbent's, but the interval runs from −4.5 to a large positive, so the claim is non-inferiority on this traffic at a large saving, not that Sonnet certifies better. claude-opus-5-5 tracks the incumbent on the point estimate at about the same cost, and does not clear the quality bound on 60 cases, so it buys neither a saving nor a proof. claude-haiku-4-5 is too weak: its control is the most sensitive of the five, so the failure is in the decisions, not the instrument.

Recovery is what makes any of these usable. The incumbent keeps the next action in 61.3% of held-out cases with distil's digests alone and 90.8% once it may call distil_expand.

What changed in the tool

The earlier published live result, 83.2% token savings at 0% decision-change, was graded by claude-opus-4-8 (2026-07-05) and is still labelled so. It was not re-graded.

What the certificate covers: the served path

distil certify grades the distil strategy, which digests only the latest, volatile tool output. A caching coding client is not served that. Its cache breakpoint sits on the newest message, so everything before it is committed prefix, and the serving adapter digests every earlier tool result. On real SWE-agent trajectories (4,616 turns) the two paths are far apart:

One request from a caching coding agent. The certified distil strategy digests only the latest tool output and saves 2.6 percent of tokens. The served path also digests every earlier tool output and saves 52.6 percent. Most served bytes were never certified.

The old certified strategy saves 2.6% of tokens. Serving saves 52.6%. The certificate was true of a slice of the traffic that carries almost none of the savings.

So there is a new strategy, distil certify --strategy served, and a served (adapter) point on the frontier. It runs the real compress_messages on the request a caching tool-use client sends and maps the result back, so the certificate covers the digests of earlier tool outputs. What it does not cover is written out in ADR 0021: a client with no or an earlier breakpoint, images, verbatim and subscription mode, and task success.

Its first result, 100 SWE-agent decision points (59 held out), graded by claude-sonnet-5-5 at low effort:

ArmNext action kept
Two identical uncompressed calls (noise floor)94.1%
Certified distil, digests alone51.7%
Certified distil, with recovery77.1%
Served adapter, digests alone59.3%
Served adapter, with recovery78.0%
This is an open risk, not a win. With recovery, the served path keeps the next action in 78.0% of held-out decisions, against 94.1% for two identical uncompressed calls. The set is small, it is one model, and a next-action match is stricter than task success because a different command can still reach the same outcome. It is also the first time the served bytes have been graded at all, and it does not look like the noise floor. The truncation control itself moves the action on only 48.3% of these decisions, so this set is also a less sensitive detector than the τ-bench one.

Whether the gap costs solved tasks is a different experiment, and it has a spec: the SWE-bench outcome eval. It runs a plain agent and a distil-served agent on SWE-bench Lite, grades with the official grader, and compares the pairs with McNemar's test against a pre-registered 5-point margin. It has been run on all of SWE-bench Lite, and it does not show that serving keeps coding tasks solved (report). The distil-served agent resolved fewer tasks than the plain agent, within the noise but on the wrong side, and it cost more, because distil digested the agent's search results and it searched and re-read again instead of expanding the digests. An earlier, smaller run that passed narrowly is superseded. Keeping shell search output verbatim (report) removes the extra steps and brings cost to parity with the plain agent, but task success is still not shown to hold, so nothing on this site claims that serving keeps coding tasks solved or makes them cheaper.

Bubble chart of the SWE-bench Lite outcome runs, one bubble per run and arm: cost per task across, share of tasks resolved up, bubble area for the number of paired tasks, coloured by plain or distil-served agent.

What is still open

Rerun it

The harness lives in benchmarks/ and calls distil's real runners. Live runs spend money; --fake oracle|null|flip checks the wiring offline first.

# fetch the pinned public trajectories (sha256-verified)
python benchmarks/model_migration_eval.py --fetch-tau-bench
python benchmarks/model_migration_eval.py --fetch-swe-agent

# the incumbent, then a candidate, on the real tau-bench cases
python benchmarks/model_migration_eval.py --cases real --flow .claude/hillclimb/decision-equivalence-real \
    --variant baseline --model claude-opus-4-8 --reps 4
python benchmarks/model_migration_eval.py --cases real --flow .claude/hillclimb/decision-equivalence-real \
    --variant v3 --model claude-sonnet-5-5 --effort low --reps 4

# the coding set, including the served arms; the free all-turns savings report
python benchmarks/model_migration_eval.py --cases coding --flow .claude/hillclimb/compression-coding \
    --variant baseline --model claude-sonnet-5-5 --effort low --reps 2
python benchmarks/model_migration_eval.py --savings-report savings.json

# status table, offline re-grade, and the summary artifact this page quotes
python benchmarks/model_migration_eval.py --status
python benchmarks/model_migration_eval.py --regrade
python benchmarks/model_migration_summary.py

Every figure on this page is recomputed by benchmarks/model_migration_summary.py from the per-case rows into benchmarks/results/model-migration/summary.json. The decision records are ADR 0020 (the default certifier) and ADR 0021 (the served strategy). The rest of the evaluation stack is on Evaluation.