compression with a quality contract

MCP compressor

distil mcp sits between any MCP client and its MCP servers and compresses what they cost you: the tool definitions resent on every request, and the tool results that pile up in context. Every level is an explicit, named mode, and every level's accuracy is decided by a pre-registered test — not by a token count.

Using Claude Code? You do not need this for definitions: Claude Code's native tool search defers unused tools better than a proxy can, and distil leaves it on (ADR 0017). This is for Codex, Cursor, Gemini CLI, opencode, Claude Desktop, Windsurf and your own agents.

Quick start

distil mcp wrap -- npx -y @modelcontextprotocol/server-filesystem ~/work   # one server
distil mcp serve --config mcp.json          # every stdio server in an mcpServers file
distil mcp install cursor                   # rewrite Cursor's config to route through distil
distil mcp install cursor --undo            # byte-exact restore
distil mcp watch                            # live view: what was compressed, per tool
distil dashboard --web                      # then open /mcp for the before/after diff

install knows cursor, claude-desktop, gemini, windsurf, opencode and codex, plus custom --path FILE for any mcpServers JSON. It keeps a byte-exact backup, writes 0600 atomically, and records what it wrote: undo restores the original bytes (and mode) when the file is untouched, and unwraps only distil's entries — keeping your later edits — when it is not. Remote (url) servers are never touched.

Levels

LevelWhat the model seesWhat is recoverableDefault
L0 losslessevery tool, with only JSON-Schema annotation keywords removed ($comment, a title that restates its name, an empty description, a no-op $schema) and whitespace normalisednothing is lost that validation reads — proven on every vendored real schemaon
L1 summarydescriptions cut to their first sentence plus up to two constraint sentences, verbatim — never generatedthe full description, via <server>_get_tool_schemaopt-in
L2 lazyone index per server; a tool's real definition appears once the model asks for itevery schema, on demandopt-in
L3 adaptiveL2, plus the tools you actually use pinned fully expanded from the first turnas L2opt-in
R resultslarge text results as distil's recoverable digestthe exact original, via <server>_expandopt-in (--results)

The default is L0 alone. L1–L3 and R passed the pre-registered live accuracy run below on claude-haiku-4-5 and are still opt-in: the protocol's default rule also needs a replication model, and that run has not been made.

No level is ever larger than L0: a server whose index or summary would not be smaller (a two-tool server, say) is served at L0, and the watch view says so. R never touches errors, images or other non-text content, results carrying structuredContent your client parses, or anything an agent may quote back byte-exact: tools named like read / view / open / cat / show / contents, and read-only tools whose description says they return file or source content. <server>_expand only returns originals that same server's results produced in the current session.

Several servers behind one proxy

distil mcp install gives every server its own proxy, so tools keep their real names. When one distil mcp serve fronts several servers, every tool, meta tool and prompt is namespaced <server>__<name> and routed through one explicit name-to-server table in config order. Server names are reduced to letters, digits and -, so one server can never mint a name that parses as another's; a name that would collide is dropped and logged, never "first claimant wins". Resources and prompts go to the server that listed them — an unknown or ambiguous one is an error, not a guess, and no kind of claim outranks another: if two servers could own a URI, by exact listing or by template, neither gets the read. Catch-all resource templates are not routed. Requests a server sends to your client (sampling, elicitation) and its instructions are labelled with the server they came from.

What it saves on definitions

On the eight MCP reference servers (78 tools, their verbatim tools/list vendored as fixtures), L0 sends 12.6% fewer model-facing definition tokens and L2 77.0% fewer at the start of a session. These are exact, deterministic counts over fixed inputs (benchmarks/results/mcp_toolbench/definitions.json, regenerated by distil mcp bench); they say nothing about accuracy, which is the next section.

Beyond wrapper-style lazy loading

Other MCP compressors fold a server's tools into one wrapper and route every call through a generic invoke_tool — the model never sees a real signature again, and every use costs an extra round trip. distil mcp does lazy loading differently:

Accuracy: pre-registered, measured

The protocol fixes everything before a live run: tool-selection accuracy and argument exact-match against the uncompressed tool list, paired, on generated tasks over the real reference-server catalog; non-inferiority by TOST and a paired bootstrap at distil's single risk budget; sample sizes from a power calculation; levels tested in a fixed sequence, one look, no rescue re-runs. distil mcp bench is its executable form. Offline it runs a scripted mock model that checks the plumbing and is never presented as evidence about a real model. With --live it needs a hard --budget-usd ceiling, which a spend meter enforces before every API call.

The first live run, on claude-haiku-4-5 (2026-09-25, $21.68 spent of an $85 ceiling), certified every level: 1,000 tool tasks per level and 630 result tasks, all eight servers exposed at once.

ArmRight toolRight tool and argumentsExtra round trips per taskBilled on this suiteVerdict
raw (uncompressed)94.4%91.2%0$2.17reference
L095.2%92.1%0$2.00certified
L195.4%95.4%0$1.90certified
L299.8%96.3%1.006$8.79certified
L398.8%95.4%0.088$1.82certified

R, on results: answers were right 99.7% of the time from raw results and 99.5% from R's digests, with the _expand tool available. Billed $3.08 raw vs $1.92 with R. Certified.

Three things to read with these numbers:

Results: benchmarks/results/mcp_toolbench/live_claude-haiku-4-5.json, per-call billing in live_calls_claude-haiku-4-5.jsonl, the analysis in the protocol.

See how it compressed

[github] level L2 (lazy)  results on  cache 3 list change(s)
  definitions 7,943 → 2,088 tok  schema fetches 2
  results 8,567 → 199 tok  calls 3  expands 1
  tool                              state              def tok  fetch  calls        result tok
  create_issue                      unlocked      290→258          1      1             37→37
  list_issues                       unlocked      371→339          1      1          4,265→81
  search_code                       unlocked      226→194          0      1          4,265→81
  add_issue_comment                 lazy          201→24           0      0                 —

Per server and per tool: definition tokens before → after, lazy / unlocked / pinned state, schema fetches, result compression, expands, how many times the list changed this session, and each level's certificate. distil dashboard --web → /mcp adds a side-by-side diff of any tool's original and sent definition, listing exactly what was dropped, and a timeline.

Everything comes from a local, content-free event log under $DISTIL_HOME/mcp/: tool names, event kinds and token sizes — never arguments, results or descriptions. Tool names stay on your machine; they are not part of the census or any telemetry.

Fail-open and transparent

Limits