compression with a quality contract

The scoreboard

Every context compressor claims a saving. This page scores them on what a user pays for: dollars per solved task, with every failed attempt's spend counted, and whether tasks still get solved. Each compressor is run by its own real, pinned package through the same harness and the same official grader. distil is one of the rows, not the referee's favourite: the same rules apply to it.

How it is measured, and why these estimators, is ADR 0024. Paired against plain on the same tasks; the interval on the difference is the harness's Wald interval with an exact McNemar test, at a pre-registered 5-point non-inferiority margin. A row that says pending run has no committed data, and nothing here estimates it. The runs that will fill it, with their costs, are in the run plan. To referee a compressor on your own traffic instead, see distil audit in the CLI reference.

SWE-bench Verified, hard tasks, effort max (official grader)

claude-sonnet-5-5 · effort max · n = 7 tasks graded in every arm · run 2026-10-06. Source: benchmarks/results/swebench-verified-hard-max. Long-horizon pilot, not a verdict: 7 tasks (the hard '1-4 hours'/'>4 hours' Verified buckets, as many as the $100 cap bought) at effort max, median 63-69 steps per task. No cost difference is shown for distil or rtk (both $-per-solved intervals cross 1). distil's lower $ per solved comes from one attempt that crashed after 4 steps on max_tokens; on the other 6 tasks distil cost 9.7% more than plain. selective and provider-cm were dropped to fit the cap (pre-registered rule).

Compressor$ per solved taskSolved (95% CI)vs plain, pts [95% CI]Total $Tokens in / outStepsPinned version
plain (no compression)$3.60506/7 · 85.7% (48.7%–97.4%)baseline$21.6352,588,959 / 825,483463—
distil$3.02676/7 · 85.7% (48.7%–97.4%)+0.0 [+0.0, +0.0] · PILOT (no verdict)$18.1646,957,534 / 658,560475distil 1.57.0
RTK$3.07737/7 · 100.0% (64.6%–100.0%)+14.3 [-11.6, +40.2] · PILOT (no verdict)$21.5455,288,254 / 773,387512rtk 0.51.0
Headroompending run
Selective Contextpending run
Anthropic context editingpending run

SWE-bench Lite (official grader)

claude-sonnet-5-5 · effort medium · n = 298 tasks graded in every arm · run 2026-10-05. Source: benchmarks/results/swebench-outcome-300-h2h. Plain and distil are the swebench-outcome-300-medium run (2026-10-03), reused unchanged, so this section supersedes that run's own table; rtk, selective and provider-cm ran 2026-10-05 on the same model and effort, provider drift between the two dates is not controlled, and provider-cm used an aggressive clearing setting rather than Anthropic's defaults.

Compressor$ per solved task (cold cache)Solved (95% CI)vs plain, pts [95% CI]Total $Tokens in / outStepsPinned version
plain (no compression)$0.0387218/298 · 73.2% (67.9%–77.9%)baseline$8.446,810,575 / 357,5621360—
distil$0.0393213/298 · 71.5% (66.1%–76.3%)-1.7 [-5.1, +1.7] · INCONCLUSIVE$8.377,081,760 / 351,38814271.56.3
RTK$0.0360†219/298 · 73.5% (68.2%–78.2%)+0.3 [-2.7, +3.3] · NON-INFERIOR$7.89†5,811,151 / 340,0321307rtk 0.51.0
Headroompending run
Selective Contextconfounded — see note210/298 · 70.5% (65.1%–75.4%)-2.7 [-6.7, +1.4] · INCONCLUSIVEconfounded — see note10,771,687 / 495,4262053selective-context 0.1.4
Anthropic context editingconfounded — see note210/298 · 70.5% (65.1%–75.4%)-2.7 [-5.6, +0.2] · INCONCLUSIVEconfounded — see note5,643,391 / 340,5891419anthropic 1.11.0

Cost confound. rtk, selective and provider-cm ran the same day with byte-identical first requests and read each other's prompt cache; plain and distil ran alone and did not. Re-pricing as a cold run keeps each task's exact prompt and output tokens and moves only the cache write/read split (benchmarks/why_rtk_wins.py, docs/research/why-rtk-wins.md). The analysis did not re-price selective, and re-priced provider-cm only on the 192 of 298 tasks without server-side edits, so neither shows a dollar figure. Success rates and verdicts are unaffected. Cold re-priced, RTK vs plain per task: -6.5% [-13.1%, +0.1%] (as billed -20.7%). † re-priced as a cold run. As billed, not comparable: RTK $0.0306 per solved, $6.69 total (190/300 rows read another arm's cache); Selective Context $0.0469 per solved, $9.86 total (167/300 rows read another arm's cache); Anthropic context editing $0.0388 per solved, $8.15 total (139/300 rows read another arm's cache).

SWE-bench Lite (official grader)

claude-sonnet-5-5 · effort low · n = 299 tasks graded in every arm · run 2026-10-02. Source: benchmarks/results/swebench-outcome-300-grepfix. Plain rows are reused from swebench-outcome-300 (same model, effort and tasks), as that run's report.md says.

Compressor$ per solved taskSolved (95% CI)vs plain, pts [95% CI]Total $Tokens in / outStepsPinned version
plain (no compression)$0.0341210/299 · 70.2% (64.8%–75.1%)baseline$7.175,114,011 / 303,2281251—
distil$0.0351204/299 · 68.2% (62.7%–73.2%)-2.0 [-5.5, +1.4] · INCONCLUSIVE$7.175,527,181 / 296,14812901.56.3
RTKpending run
Headroompending run
Selective Contextpending run
Anthropic context editingpending run

Terminal-Bench 2.1 pilot (cost_truth, neutral meter)

claude-sonnet-5 · n = 10 tasks graded in every arm · run 2026-09-25. Source: benchmarks/results/cost_truth/pilot-20260925-145427. Pilot, not a comparison (protocol section 7): 18 paired attempts per arm after 10 infra_error runs were excluded, run under emulated x86_64, and billed while distil priced claude-sonnet-5 at $3/$15 instead of the official $2/$10, so absolute dollars are 1.5x too high and ratios are unaffected. analysis.json carries no passing verdict for any arm.

Compressor$ per solved taskSolved (95% CI)vs plain, pts [95% CI]Total $Tokens in / outStepsPinned version
control (no compression)$0.808814/18 · 77.8%baseline$11.32———
distil$0.982314/18 · 77.8%+0.0 [-33.3, +35.0] · pilot, not shown$13.75——1.54.0
RTK$0.598315/18 · 83.3%+5.6 [-17.6, +25.0] · pilot, not shown$8.97——0.50.0
Headroom$0.781213/18 · 72.2%-5.6 [-29.4, +11.8] · pilot, not shown$10.16——0.38.0
Selective Contextnot in this harness
Anthropic context editingnot in this harness

How to read it