Model migration
A distil certificate is graded by a model. Change that model and you have changed what the certificate means, so distil does not change it on a hunch. This page is the eval that decided the default live certifier, the gates it had to clear, what it found about the served path while doing it, and how to rerun all of it. One result is good news and one is an open risk; both are here.
The question
Moving the grader from an incumbent model to a newer one is a migration, and a migration has one test: given the same contexts, compressed and not, does the new model reach the same verdicts? The verdict distil cares about is narrow. Is the agent's next action unchanged by compression, when the model is allowed to recover a digest with distil_expand? That is the metric distil is deployed on, and the one this eval treats as the headline.
The eval asks it twice, about two different things. First, which model should grade (the certifier result). Second, does the path distil actually serves to a caching client stay inside the same contract (the served result). The second question only arose because the first eval's harness would not let it stay hidden.
Method
Each case is one decision point from a real trajectory: a τ-bench turn, or a SWE-agent step. Per case and repetition, distil's own runners answer the same question under six arms.
Data. Public τ-bench historical trajectories (airline and retail, GPT-4o and Claude Sonnet agents), SWE-agent GPT-4o trajectories from the SWE-bench Lite leaderboard submissions, and SWE-bench Lite itself for the task-outcome runs. Every download is pinned and hash-verified; the full table, with files, selection seeds and pins, is in EVALUATION.md §6.9.
- Full, twice. Uncompressed context, two independent calls. Their agreement is the certifier's own noise: if two identical calls disagree, a disagreement between full and compressed proves nothing.
- Certified distil, and with recovery. The cache-aware strategy
distil certifyhas always graded, alone and with the recovery loop that lets the model calldistil_expand. - Served adapter, and with recovery. The real
distil.adapters.anthropic.compress_messageson the request a caching tool-use client sends. See below. - Truncation control. Tool output cut short. It should change the action. If it does not, the grader could not have seen a harm, and an unchanged score means nothing.
Cases are paired: each candidate grades exactly the cases the incumbent did, each case averaged over its repetitions, and the difference carries a 95% interval from the case-level differences. Cost is the provider's own usage at list price over every call. Candidates are compared on held-out cases, set aside before any candidate was run. A harness hash gates every run, so a change to the runners or prompts has to be approved before its numbers can be mixed with older ones.
Two lessons came out of building it, and both would have ruined a result silently:
- A distil-wrapped shell points the client at distil.
distil wrapexportsANTHROPIC_BASE_URLto the local proxy, so an eval client that honours it grades distil with distil. The eval client pinsapi.anthropic.com. - The bundled corpus is for the offline oracle. It plants
DECISION:markers that the deterministic runner reads and that a live model reads too, which hands it the answer. Live grading needs marker-free real traces, and live renders now strip the marker lines.
Which model certifies: the result
The real-traffic set is 100 τ-bench decision points, 60 held out, each run 4 times. The incumbent was claude-opus-4-8. Candidates were the newer claude-opus-5-5 (default effort, and low), claude-sonnet-5-5 at low effort, and claude-haiku-4-5.
| Certifier | Next action kept, with recovery | Paired change, lower bound (pts) | Distil alone, no recovery | Two full calls agree | Control diverges | Cost per case |
|---|---|---|---|---|---|---|
claude-opus-4-8 (incumbent) | 90.8% | reference | 61.3% | 92.5% | 89.2% | $0.1458 |
claude-opus-5-5 | 90.8% | +0.0, −10.0 | 73.3% | 93.3% | 73.3% | $0.1336 |
claude-opus-5-5, effort low | 90.0% | −0.8, −9.8 | 75.0% | 94.2% | 75.8% | $0.1256 |
claude-sonnet-5-5, effort low | 94.6% | +3.8, −4.5 | 66.7% | 99.2% | 79.6% | $0.0619 |
claude-haiku-4-5 | 85.0% | −5.8, −15.8 | 44.2% | 96.7% | 93.3% | $0.0248 |
The default is now claude-sonnet-5-5 at effort=low. It was accepted under five gates written down before the run, and it passed all of them:
- Quality. The lower bound of the paired change in the deployed metric is at least −5 points. It is −4.5.
- Self-consistency. Two uncompressed calls agree at least as often as the incumbent's do. They agree more often, 99.2% against 92.5%.
- The control still bites. Truncation must still diverge, at no less than 0.8× the incumbent's rate. It diverges on 79.6% of held-out cases against 89.2%.
- Cost. At least 15% cheaper per case. It is 57.5% cheaper, $0.0619 against $0.1458.
- Mechanism. A difference has to be traced to something in the traces, not accepted as an unexplained delta.
Read it for what it is. The point estimate is above the incumbent's, but the interval runs from −4.5 to a large positive, so the claim is non-inferiority on this traffic at a large saving, not that Sonnet certifies better. claude-opus-5-5 tracks the incumbent on the point estimate at about the same cost, and does not clear the quality bound on 60 cases, so it buys neither a saving nor a proof. claude-haiku-4-5 is too weak: its control is the most sensitive of the five, so the failure is in the decisions, not the instrument.
Recovery is what makes any of these usable. The incumbent keeps the next action in 61.3% of held-out cases with distil's digests alone and 90.8% once it may call distil_expand.
What changed in the tool
AnthropicRunnerdefaults to the pair above, and the nightly live gate moved fromclaude-haiku-4-5to it.- A new
--effortflag oncertify,eval,benchmark,frontierandconformal. An explicit value wins,nonesends nooutput_config, and unset meanslowfor the default model and nothing for any other, because some models (claude-haiku-4-5) reject one. - The newer models (
claude-opus-5-5,claude-sonnet-5-5,claude-fable-5-1) reject a forcedtool_choicewith a 400.AnthropicRunnernow falls back totool_choice=autowith the same strict decision tool; before,--runner anthropicexited on them. - The recovery loop commits its final decision through the runner's structured decision tool. The free-text format difference alone had flipped about one action in five, which read as compression harm and was a grading artefact.
The earlier published live result, 83.2% token savings at 0% decision-change, was graded by claude-opus-4-8 (2026-07-05) and is still labelled so. It was not re-graded.
What the certificate covers: the served path
distil certify grades the distil strategy, which digests only the latest, volatile tool output. A caching coding client is not served that. Its cache breakpoint sits on the newest message, so everything before it is committed prefix, and the serving adapter digests every earlier tool result. On real SWE-agent trajectories (4,616 turns) the two paths are far apart:
The old certified strategy saves 2.6% of tokens. Serving saves 52.6%. The certificate was true of a slice of the traffic that carries almost none of the savings.
So there is a new strategy, distil certify --strategy served, and a served (adapter) point on the frontier. It runs the real compress_messages on the request a caching tool-use client sends and maps the result back, so the certificate covers the digests of earlier tool outputs. What it does not cover is written out in ADR 0021: a client with no or an earlier breakpoint, images, verbatim and subscription mode, and task success.
Its first result, 100 SWE-agent decision points (59 held out), graded by claude-sonnet-5-5 at low effort:
| Arm | Next action kept |
|---|---|
| Two identical uncompressed calls (noise floor) | 94.1% |
Certified distil, digests alone | 51.7% |
Certified distil, with recovery | 77.1% |
| Served adapter, digests alone | 59.3% |
| Served adapter, with recovery | 78.0% |
Whether the gap costs solved tasks is a different experiment, and it has a spec: the SWE-bench outcome eval. It runs a plain agent and a distil-served agent on SWE-bench Lite, grades with the official grader, and compares the pairs with McNemar's test against a pre-registered 5-point margin. It has been run on all of SWE-bench Lite, and it does not show that serving keeps coding tasks solved (report). The distil-served agent resolved fewer tasks than the plain agent, within the noise but on the wrong side, and it cost more, because distil digested the agent's search results and it searched and re-read again instead of expanding the digests. An earlier, smaller run that passed narrowly is superseded. Keeping shell search output verbatim (report) removes the extra steps and brings cost to parity with the plain agent, but task success is still not shown to hold, so nothing on this site claims that serving keeps coding tasks solved or makes them cheaper.
What is still open
- The served path's decision-level gap on coding traffic, and whether it costs tasks (the outcome eval, unrun).
- The certifier choice rests on one task family and 60 held-out cases. The intervals are wide. A replication model and a coding-traffic migration run have not been done.
- Clients with no cache breakpoint, or one earlier in the history, are not exercised by the served strategy.
claude-fable-5-1is handled by the runner but was not measured as a certifier.
Rerun it
The harness lives in benchmarks/ and calls distil's real runners. Live runs spend money; --fake oracle|null|flip checks the wiring offline first.
# fetch the pinned public trajectories (sha256-verified) python benchmarks/model_migration_eval.py --fetch-tau-bench python benchmarks/model_migration_eval.py --fetch-swe-agent # the incumbent, then a candidate, on the real tau-bench cases python benchmarks/model_migration_eval.py --cases real --flow .claude/hillclimb/decision-equivalence-real \ --variant baseline --model claude-opus-4-8 --reps 4 python benchmarks/model_migration_eval.py --cases real --flow .claude/hillclimb/decision-equivalence-real \ --variant v3 --model claude-sonnet-5-5 --effort low --reps 4 # the coding set, including the served arms; the free all-turns savings report python benchmarks/model_migration_eval.py --cases coding --flow .claude/hillclimb/compression-coding \ --variant baseline --model claude-sonnet-5-5 --effort low --reps 2 python benchmarks/model_migration_eval.py --savings-report savings.json # status table, offline re-grade, and the summary artifact this page quotes python benchmarks/model_migration_eval.py --status python benchmarks/model_migration_eval.py --regrade python benchmarks/model_migration_summary.py
Every figure on this page is recomputed by benchmarks/model_migration_summary.py from the per-case rows into benchmarks/results/model-migration/summary.json. The decision records are ADR 0020 (the default certifier) and ADR 0021 (the served strategy). The rest of the evaluation stack is on Evaluation.