MCP compressor
distil mcp sits between any MCP client and its MCP servers and compresses what they cost you: the tool definitions resent on every request, and the tool results that pile up in context. Every level is an explicit, named mode, and every level's accuracy is decided by a pre-registered test — not by a token count.
Using Claude Code? You do not need this for definitions: Claude Code's native tool search defers unused tools better than a proxy can, and distil leaves it on (ADR 0017). This is for Codex, Cursor, Gemini CLI, opencode, Claude Desktop, Windsurf and your own agents.
Quick start
distil mcp wrap -- npx -y @modelcontextprotocol/server-filesystem ~/work # one server
distil mcp serve --config mcp.json # every stdio server in an mcpServers file
distil mcp install cursor # rewrite Cursor's config to route through distil
distil mcp install cursor --undo # byte-exact restore
distil mcp watch # live view: what was compressed, per tool
distil dashboard --web # then open /mcp for the before/after diff
install knows cursor, claude-desktop, gemini, windsurf, opencode and codex, plus custom --path FILE for any mcpServers JSON. It keeps a byte-exact backup, writes 0600 atomically, and records what it wrote: undo restores the original bytes (and mode) when the file is untouched, and unwraps only distil's entries — keeping your later edits — when it is not. Remote (url) servers are never touched.
Levels
| Level | What the model sees | What is recoverable | Default |
|---|---|---|---|
| L0 lossless | every tool, with only JSON-Schema annotation keywords removed ($comment, a title that restates its name, an empty description, a no-op $schema) and whitespace normalised | nothing is lost that validation reads — proven on every vendored real schema | on |
| L1 summary | descriptions cut to their first sentence plus up to two constraint sentences, verbatim — never generated | the full description, via <server>_get_tool_schema | opt-in |
| L2 lazy | one index per server; a tool's real definition appears once the model asks for it | every schema, on demand | opt-in |
| L3 adaptive | L2, plus the tools you actually use pinned fully expanded from the first turn | as L2 | opt-in |
| R results | large text results as distil's recoverable digest | the exact original, via <server>_expand | opt-in (--results) |
The default is L0 alone. L1–L3 and R passed the pre-registered live accuracy run below on claude-haiku-4-5 and are still opt-in: the protocol's default rule also needs a replication model, and that run has not been made.
No level is ever larger than L0: a server whose index or summary would not be smaller (a two-tool server, say) is served at L0, and the watch view says so. R never touches errors, images or other non-text content, results carrying structuredContent your client parses, or anything an agent may quote back byte-exact: tools named like read / view / open / cat / show / contents, and read-only tools whose description says they return file or source content. <server>_expand only returns originals that same server's results produced in the current session.
Several servers behind one proxy
distil mcp install gives every server its own proxy, so tools keep their real names. When one distil mcp serve fronts several servers, every tool, meta tool and prompt is namespaced <server>__<name> and routed through one explicit name-to-server table in config order. Server names are reduced to letters, digits and -, so one server can never mint a name that parses as another's; a name that would collide is dropped and logged, never "first claimant wins". Resources and prompts go to the server that listed them — an unknown or ambiguous one is an error, not a guess, and no kind of claim outranks another: if two servers could own a URI, by exact listing or by template, neither gets the read. Catch-all resource templates are not routed. Requests a server sends to your client (sampling, elicitation) and its instructions are labelled with the server they came from.
What it saves on definitions
On the eight MCP reference servers (78 tools, their verbatim tools/list vendored as fixtures), L0 sends 12.6% fewer model-facing definition tokens and L2 77.0% fewer at the start of a session. These are exact, deterministic counts over fixed inputs (benchmarks/results/mcp_toolbench/definitions.json, regenerated by distil mcp bench); they say nothing about accuracy, which is the next section.
Beyond wrapper-style lazy loading
Other MCP compressors fold a server's tools into one wrapper and route every call through a generic invoke_tool — the model never sees a real signature again, and every use costs an extra round trip. distil mcp does lazy loading differently:
- The real tool comes back. Fetching a schema adds the tool to the list under its real name with its real
inputSchema, and tells the client (notifications/tools/list_changed). The next call is a normal, typed tool call.<server>_invoke_toolremains only as the fallback for clients that never refresh their list. - Cache-stable. The index is one byte-stable description; unlocked tools are appended at the end and never removed, and the set is persisted per session, so a cached prefix survives an unlock.
- Learns what you use. L3 pins the tools a server is actually called for — tool names and counts only, on this machine — and fixes the pins for the session.
- Compresses results too, recoverably, for every client.
- Publishes accuracy, not just tokens — below.
Accuracy: pre-registered, measured
The protocol fixes everything before a live run: tool-selection accuracy and argument exact-match against the uncompressed tool list, paired, on generated tasks over the real reference-server catalog; non-inferiority by TOST and a paired bootstrap at distil's single risk budget; sample sizes from a power calculation; levels tested in a fixed sequence, one look, no rescue re-runs. distil mcp bench is its executable form. Offline it runs a scripted mock model that checks the plumbing and is never presented as evidence about a real model. With --live it needs a hard --budget-usd ceiling, which a spend meter enforces before every API call.
The first live run, on claude-haiku-4-5 (2026-09-25, $21.68 spent of an $85 ceiling), certified every level: 1,000 tool tasks per level and 630 result tasks, all eight servers exposed at once.
| Arm | Right tool | Right tool and arguments | Extra round trips per task | Billed on this suite | Verdict |
|---|---|---|---|---|---|
| raw (uncompressed) | 94.4% | 91.2% | 0 | $2.17 | reference |
| L0 | 95.2% | 92.1% | 0 | $2.00 | certified |
| L1 | 95.4% | 95.4% | 0 | $1.90 | certified |
| L2 | 99.8% | 96.3% | 1.006 | $8.79 | certified |
| L3 | 98.8% | 95.4% | 0.088 | $1.82 | certified |
R, on results: answers were right 99.7% of the time from raw results and 99.5% from R's digests, with the _expand tool available. Billed $3.08 raw vs $1.92 with R. Certified.
Three things to read with these numbers:
- Compressed levels scored higher than raw. The protocol tests non-inferiority, not superiority, so that is an observation, not a claim.
- L2 cost more, not less. Its short index was never written to the prompt cache on this model, and every task took a second round trip, so in this benchmark (one fresh session per task) it billed $8.79 against raw's $2.17. Its value is context-window headroom. L3 was cheaper than raw only because its pins had already been learned; a fresh L3 install starts as L2.
- Scope. One model, the eight reference servers, synthetic single-call tasks. No replication model was run, so the default stays L0 and L1–L3 and R stay opt-in.
distil mcp watchand the/mcpdashboard show them as certified.
Results: benchmarks/results/mcp_toolbench/live_claude-haiku-4-5.json, per-call billing in live_calls_claude-haiku-4-5.jsonl, the analysis in the protocol.
See how it compressed
[github] level L2 (lazy) results on cache 3 list change(s)
definitions 7,943 → 2,088 tok schema fetches 2
results 8,567 → 199 tok calls 3 expands 1
tool state def tok fetch calls result tok
create_issue unlocked 290→258 1 1 37→37
list_issues unlocked 371→339 1 1 4,265→81
search_code unlocked 226→194 0 1 4,265→81
add_issue_comment lazy 201→24 0 0 —
Per server and per tool: definition tokens before → after, lazy / unlocked / pinned state, schema fetches, result compression, expands, how many times the list changed this session, and each level's certificate. distil dashboard --web → /mcp adds a side-by-side diff of any tool's original and sent definition, listing exactly what was dropped, and a timeline.
Everything comes from a local, content-free event log under $DISTIL_HOME/mcp/: tool names, event kinds and token sizes — never arguments, results or descriptions. Tool names stay on your machine; they are not part of the census or any telemetry.
Fail-open and transparent
- Only
tools/listandtools/callare rewritten. Resources, prompts, completion, logging, notifications and server-to-client requests (roots/list, sampling) pass through, with request ids preserved so cancellation and progress keep working. - A compression error returns the server's own answer. If the proxy cannot start at all,
wrapruns the server directly. - The server runs with the command, environment and working directory it was configured with; its stderr goes wherever your client already collects it.
Limits
- stdio servers only. Remote (streamable-HTTP / SSE) servers are skipped by
serveand left alone byinstall. - Rewriting Codex's
config.tomlneeds Python 3.11+; the patch is verified by re-parsing, and anything that would not round-trip is refused rather than written. - A client that ignores
list_changedgets L2/L3 through the invoke fallback — typed calls are the part it misses.