The scoreboard
Every context compressor claims a saving. This page scores them on what a user pays for: dollars per solved task, with every failed attempt's spend counted, and whether tasks still get solved. Each compressor is run by its own real, pinned package through the same harness and the same official grader. distil is one of the rows, not the referee's favourite: the same rules apply to it.
How it is measured, and why these estimators, is ADR 0024. Paired against plain on the same tasks; the interval on the difference is the harness's Wald interval with an exact McNemar test, at a pre-registered 5-point non-inferiority margin. A row that says pending run has no committed data, and nothing here estimates it. The runs that will fill it, with their costs, are in the run plan. To referee a compressor on your own traffic instead, see distil audit in the CLI reference.
SWE-bench Verified, hard tasks, effort max (official grader)
claude-sonnet-5-5 · effort max · n = 7 tasks graded in every arm · run 2026-10-06. Source: benchmarks/results/swebench-verified-hard-max. Long-horizon pilot, not a verdict: 7 tasks (the hard '1-4 hours'/'>4 hours' Verified buckets, as many as the $100 cap bought) at effort max, median 63-69 steps per task. No cost difference is shown for distil or rtk (both $-per-solved intervals cross 1). distil's lower $ per solved comes from one attempt that crashed after 4 steps on max_tokens; on the other 6 tasks distil cost 9.7% more than plain. selective and provider-cm were dropped to fit the cap (pre-registered rule).
| Compressor | $ per solved task | Solved (95% CI) | vs plain, pts [95% CI] | Total $ | Tokens in / out | Steps | Pinned version |
|---|---|---|---|---|---|---|---|
| plain (no compression) | $3.6050 | 6/7 · 85.7% (48.7%–97.4%) | baseline | $21.63 | 52,588,959 / 825,483 | 463 | — |
| distil | $3.0267 | 6/7 · 85.7% (48.7%–97.4%) | +0.0 [+0.0, +0.0] · PILOT (no verdict) | $18.16 | 46,957,534 / 658,560 | 475 | distil 1.57.0 |
| RTK | $3.0773 | 7/7 · 100.0% (64.6%–100.0%) | +14.3 [-11.6, +40.2] · PILOT (no verdict) | $21.54 | 55,288,254 / 773,387 | 512 | rtk 0.51.0 |
| Headroom | pending run | ||||||
| Selective Context | pending run | ||||||
| Anthropic context editing | pending run | ||||||
SWE-bench Lite (official grader)
claude-sonnet-5-5 · effort medium · n = 298 tasks graded in every arm · run 2026-10-05. Source: benchmarks/results/swebench-outcome-300-h2h. Plain and distil are the swebench-outcome-300-medium run (2026-10-03), reused unchanged, so this section supersedes that run's own table; rtk, selective and provider-cm ran 2026-10-05 on the same model and effort, provider drift between the two dates is not controlled, and provider-cm used an aggressive clearing setting rather than Anthropic's defaults.
| Compressor | $ per solved task (cold cache) | Solved (95% CI) | vs plain, pts [95% CI] | Total $ | Tokens in / out | Steps | Pinned version |
|---|---|---|---|---|---|---|---|
| plain (no compression) | $0.0387 | 218/298 · 73.2% (67.9%–77.9%) | baseline | $8.44 | 6,810,575 / 357,562 | 1360 | — |
| distil | $0.0393 | 213/298 · 71.5% (66.1%–76.3%) | -1.7 [-5.1, +1.7] · INCONCLUSIVE | $8.37 | 7,081,760 / 351,388 | 1427 | 1.56.3 |
| RTK | $0.0360† | 219/298 · 73.5% (68.2%–78.2%) | +0.3 [-2.7, +3.3] · NON-INFERIOR | $7.89† | 5,811,151 / 340,032 | 1307 | rtk 0.51.0 |
| Headroom | pending run | ||||||
| Selective Context | confounded — see note | 210/298 · 70.5% (65.1%–75.4%) | -2.7 [-6.7, +1.4] · INCONCLUSIVE | confounded — see note | 10,771,687 / 495,426 | 2053 | selective-context 0.1.4 |
| Anthropic context editing | confounded — see note | 210/298 · 70.5% (65.1%–75.4%) | -2.7 [-5.6, +0.2] · INCONCLUSIVE | confounded — see note | 5,643,391 / 340,589 | 1419 | anthropic 1.11.0 |
Cost confound. rtk, selective and provider-cm ran the same day with byte-identical first requests and read each other's prompt cache; plain and distil ran alone and did not. Re-pricing as a cold run keeps each task's exact prompt and output tokens and moves only the cache write/read split (benchmarks/why_rtk_wins.py, docs/research/why-rtk-wins.md). The analysis did not re-price selective, and re-priced provider-cm only on the 192 of 298 tasks without server-side edits, so neither shows a dollar figure. Success rates and verdicts are unaffected. Cold re-priced, RTK vs plain per task: -6.5% [-13.1%, +0.1%] (as billed -20.7%). † re-priced as a cold run. As billed, not comparable: RTK $0.0306 per solved, $6.69 total (190/300 rows read another arm's cache); Selective Context $0.0469 per solved, $9.86 total (167/300 rows read another arm's cache); Anthropic context editing $0.0388 per solved, $8.15 total (139/300 rows read another arm's cache).
SWE-bench Lite (official grader)
claude-sonnet-5-5 · effort low · n = 299 tasks graded in every arm · run 2026-10-02. Source: benchmarks/results/swebench-outcome-300-grepfix. Plain rows are reused from swebench-outcome-300 (same model, effort and tasks), as that run's report.md says.
| Compressor | $ per solved task | Solved (95% CI) | vs plain, pts [95% CI] | Total $ | Tokens in / out | Steps | Pinned version |
|---|---|---|---|---|---|---|---|
| plain (no compression) | $0.0341 | 210/299 · 70.2% (64.8%–75.1%) | baseline | $7.17 | 5,114,011 / 303,228 | 1251 | — |
| distil | $0.0351 | 204/299 · 68.2% (62.7%–73.2%) | -2.0 [-5.5, +1.4] · INCONCLUSIVE | $7.17 | 5,527,181 / 296,148 | 1290 | 1.56.3 |
| RTK | pending run | ||||||
| Headroom | pending run | ||||||
| Selective Context | pending run | ||||||
| Anthropic context editing | pending run | ||||||
Terminal-Bench 2.1 pilot (cost_truth, neutral meter)
claude-sonnet-5 · n = 10 tasks graded in every arm · run 2026-09-25. Source: benchmarks/results/cost_truth/pilot-20260925-145427. Pilot, not a comparison (protocol section 7): 18 paired attempts per arm after 10 infra_error runs were excluded, run under emulated x86_64, and billed while distil priced claude-sonnet-5 at $3/$15 instead of the official $2/$10, so absolute dollars are 1.5x too high and ratios are unaffected. analysis.json carries no passing verdict for any arm.
| Compressor | $ per solved task | Solved (95% CI) | vs plain, pts [95% CI] | Total $ | Tokens in / out | Steps | Pinned version |
|---|---|---|---|---|---|---|---|
| control (no compression) | $0.8088 | 14/18 · 77.8% | baseline | $11.32 | — | — | — |
| distil | $0.9823 | 14/18 · 77.8% | +0.0 [-33.3, +35.0] · pilot, not shown | $13.75 | — | — | 1.54.0 |
| RTK | $0.5983 | 15/18 · 83.3% | +5.6 [-17.6, +25.0] · pilot, not shown | $8.97 | — | — | 0.50.0 |
| Headroom | $0.7812 | 13/18 · 72.2% | -5.6 [-29.4, +11.8] · pilot, not shown | $10.16 | — | — | 0.38.0 |
| Selective Context | not in this harness | ||||||
| Anthropic context editing | not in this harness | ||||||
How to read it
- $ per solved task is the arm's total spend on the tasks graded in every arm, divided by the tasks it solved. A cheaper arm that solves fewer tasks can cost more per solved task.
- Cold cache. Each arm's dollars must be its own. The builder checks every row: a run that started cold writes at least its largest prompt, so
cache_write + input < cache_read / (steps − 1)means it read a prompt cache another arm wrote. Such an arm shows a cold re-pricing, or no dollar figure, never the billed one. - vs plain is the paired difference in tasks solved, in points, with its verdict at the 5-point margin. PILOT, INCONCLUSIVE and INFERIOR mean what they say; none of them is a pass.
- Short tasks. On these runs the median task took a handful of steps, so contexts stayed short. A compressor built for long sessions is not exercised much here; the Terminal-Bench table runs longer agent sessions.
- Every figure is regenerated by
scripts/build_scoreboard.pyfrom the run directories, and the machine-readable copy is scoreboard.json.