compression with a quality contract

Integrations — one proxy, every SDK

Start distil proxy once and point any SDK's baseURL at it. No library changes, no monkey-patching — compression happens at the network layer.

Cross-SDK proxy diagram

How it works

The proxy (distil proxy, default http://127.0.0.1:8788) is a local HTTP server. It intercepts the three compressible paths across all major LLM APIs:

All other paths and HTTP verbs pass through unchanged. Your API key travels in the request headers exactly as normal — the proxy never logs or stores it.

Zero SDK changes beyond baseURL. The proxy is transparent: the SDK sees the same wire format it always expects. Auth headers, streaming, tool use, and all other features work as-is.

SDK integration matrix

SDK / Framework Language Setting Value Example
Anthropic Python SDK Python base_url= http://127.0.0.1:8788 python_anthropic.py
Anthropic TypeScript SDK (@anthropic-ai/sdk) TypeScript baseURL in new Anthropic({…}) http://127.0.0.1:8788 js_anthropic.ts
Claude Agent SDK / claude -p (headless) Python / TS / CLI distil wrap -- <cmd> or ANTHROPIC_BASE_URL http://127.0.0.1:8788 python_claude_agent_sdk.py
OpenAI Python SDK Python base_url= http://127.0.0.1:8788/v1 python_openai.py
LiteLLM Python api_base= http://127.0.0.1:8788 python_litellm.py
Cursor (agent/chat panel only) IDE Settings → Override OpenAI Base URL http://127.0.0.1:8788/v1 below
CrewAI Python base_url= in LLM({…}) http://127.0.0.1:8788/v1 below
Agno Python base_url= in OpenAILike({…}) http://127.0.0.1:8788/v1 below
Strands Agents Python client_args={"base_url": …} in OpenAIModel({…}) http://127.0.0.1:8788/v1 below
Microsoft AutoGen Python base_url= in OpenAIChatCompletionClient({…}) http://127.0.0.1:8788/v1 below
LlamaIndex Python api_base= in OpenAI({…}) http://127.0.0.1:8788/v1 below
Vercel AI SDK (@ai-sdk/anthropic) TypeScript baseURL in createAnthropic({…}) http://127.0.0.1:8788 js_vercel_ai_sdk.ts
LangChain.js (@langchain/anthropic) TypeScript anthropicApiUrl in ChatAnthropic({…}) http://127.0.0.1:8788 js_langchain.ts
Google Gemini REST (google-generativeai) Python / curl api_endpoint in client_options http://127.0.0.1:8788 (upstream: https://generativelanguage.googleapis.com) python_gemini.py

Agents distil wrap routes for you

No configuration and no code change: distil wrap -- <agent> starts a proxy, points the agent at it, and restores everything on exit. Most read an environment variable; a few have no such contract and route through a config file wrap manages for that one session. This table and the next are generated from distil/targets.py — distil wrap --list prints the same thing in your terminal.

AgentCommandMechanismRouting knobWire shape
aiderdistil wrap -- aiderenvironment variableOPENAI_API_BASEOpenAI Chat Completions
Claude Codedistil wrap -- claudeenvironment variableANTHROPIC_BASE_URLAnthropic Messages
Codex CLIdistil wrap -- codexenvironment variableOPENAI_BASE_URLOpenAI Responses
Gemini CLIdistil wrap -- geminienvironment variableGOOGLE_GEMINI_BASE_URLGemini generateContent
GitHub Copilot CLIdistil wrap -- copilotenvironment variableCOPILOT_PROVIDER_BASE_URLAnthropic Messages
goosedistil wrap -- gooseenvironment variableOPENAI_HOSTOpenAI Chat Completions
Grok CLIdistil wrap -- grokenvironment variableGROK_MODELS_BASE_URLOpenAI Chat Completions
Kilo Code CLIdistil wrap -- kiloenvironment variableKILO_CONFIG_CONTENTAnthropic Messages or OpenAI Chat Completions
Kimi CLIdistil wrap -- kimienvironment variableKIMI_BASE_URLOpenAI Chat Completions
Mistral Vibedistil wrap -- vibeenvironment variableVIBE_PROVIDERSOpenAI Chat Completions
OpenCodedistil wrap -- opencodeenvironment variableOPENAI_BASE_URLOpenAI Responses
OpenHandsdistil wrap -- openhandsenvironment variableLLM_BASE_URLOpenAI Chat Completions
Qwen Codedistil wrap -- qwenenvironment variableOPENAI_BASE_URLOpenAI Chat Completions
Clinedistil wrap -- clineconfig fileproviders.json → providers.<id>.settings.baseUrlAnthropic Messages or OpenAI Chat Completions
Continuedistil wrap -- cnconfig fileconfig.yaml → models[].apiBase (via --config)Anthropic Messages or OpenAI Chat Completions
Crushdistil wrap -- crushconfig filecrush.json → providers.<id>.base_urlAnthropic Messages or OpenAI Chat Completions
Factory Droiddistil wrap -- droidconfig filesettings.local.json → customModels[].baseUrlOpenAI Chat Completions
Oh My Pidistil wrap -- ompconfig filemodels.yml → baseUrlAnthropic Messages or OpenAI Chat Completions

Agents it cannot reach — and what was checked

Each of these was read against its own primary documentation on the date shown. Where a base-URL setting exists, point it at the URL distil proxy prints (http://127.0.0.1:8788/v1) and verify with distil doctor; where it says — there is no such setting to point. Full reasoning: IDE-AGENTS.md.

AgentIts own base-URL settingWhy wrap cannotVerified
Amp— (HTTP_PROXY/HTTPS_PROXY only)re-checked: the CLI settings reference still has no base-URL key; amp.url belongs to the VS Code extension, not the CLIverified 2026-09-16
Augment (auggie)—AUGMENT_SESSION_AUTH carries the session token; no base-URL variable or config key is documentedverified 2026-09-16
Continue (VS Code extension)~/.continue/config.yaml → models[].apiBaseapiBase 'can be used to override the default API base', but the extension is started by the editor — no argv to wrap, and the file is editor-wide rather than per-session. The Continue CLI is a different tool and `distil wrap -- cn` does reach itverified 2026-09-16
Cursor CLI— (HTTP_PROXY/HTTPS_PROXY only)cli-config.json publishes no base-URL field; the only network knob is a whole-process HTTP proxy, not a per-request base URL. Its binary is `agent`, a name too generic for distil to claimverified 2026-09-16
Google Antigravity—models are plan-selected from a fixed list; no BYOK and no endpoint override is documentedverified 2026-09-16
JetBrains Juniemodel profile → baseUrldocs render client-side and return nothing over a plain fetch; the config shape could not be verifiedverified 2026-09-16
OpenClaw~/.openclaw/openclaw.json → models.providers.<id>.baseUrlthe knob is verified, but OpenClaw's own README puts the model connection in a Gateway the CLI merely 'connects to' — one local control plane shared with Discord/WhatsApp/Slack channels and, on a team install, other people. A session-scoped patch would reconfigure a daemon serving them, and restore it mid-flight when one terminal exitsverified 2026-09-16
Roo CodeAPI Provider → "OpenAI Compatible" → Base URLa VS Code extension with no CLI, whose configuration profiles live in VS Code's own Secret Storage ('stored securely in VSCode's Secret Storage and never exposed in plain text') — there is no config file to patchverified 2026-09-16
Snowflake Cortex Code—no published base-URL overrideverified 2026-09-16
Sourcegraph Codysite config → modelConfiguration.providerOverrides (server-side)the override is an admin setting on the Sourcegraph instance, not on the client — a distil gateway in front of that instance is the fit, not wrapverified 2026-09-16
Tabnine—clients point at a Tabnine server, not at an LLM endpoint; the CLI docs publish no model base-URL overrideverified 2026-09-16
TraeSettings → Models → custom modelcustom models exist, but every docs path returns the same client-rendered shell over a plain fetch — the config shape could not be verified against an authoritative sourceverified 2026-09-16
VS Code Copilot Chat (extension)chatLanguageModels.json → vendor customendpoint → models[].urlBYOK 'Custom Endpoint' takes a full per-model URL (Chat Completions, Responses or Anthropic Messages), so point it at a running distil proxy — `distil setup --vscode` prints the entry. The extension is editor-launched, so there is nothing to wrap; chat only (inline completions, embeddings and semantic search stay on GitHub), and a Business/Enterprise admin can disable BYOKverified 2026-09-25
WarpSettings → custom inference endpoint (public HTTPS URL only)Warp DOES publish an endpoint override now (the older 'no override at all' note was stale) — but the agent harness runs on Warp's servers and the docs reject localhost and private addresses, so a local distil proxy cannot be the targetverified 2026-09-16
WindsurfSettings → Cascade → custom endpointBYOK accepts a provider API KEY only (Claude 4 family), with no endpoint field in the documented flowverified 2026-09-16
ZCode (z.ai)Settings → Providers → Base URL (Anthropic or OpenAI protocol)z.ai's own page calls it an Agentic Development Environment, a desktop app with no CLI — the Base URL field is verified and does take a local proxy, but there is no process for wrap to launch or scope a config change toverified 2026-09-16
Zed agentsettings.json → language_models.anthropic_compatible.<name>.api_url (or openai_compatible)no single key redirects it: the BUILT-IN anthropic provider documents only available_models, never an api_url, so routing means ADDING an anthropic_compatible provider the user must then pick by hand in the model dropdown — and Zed's own Agent Settings page writes settings.json while it runs, so a session-scoped patch would be racing the editor for the fileverified 2026-09-16

Snippets

Anthropic Python SDK

import anthropic

client = anthropic.Anthropic(
    api_key="sk-ant-…",
    base_url="http://127.0.0.1:8788",
)
response = client.messages.create(
    model="claude-opus-4-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello!"}],
)

OpenAI Python SDK

import openai

client = openai.OpenAI(
    api_key="sk-…",
    base_url="http://127.0.0.1:8788/v1",
)
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}],
)

LiteLLM

import litellm

response = litellm.completion(
    model="claude-opus-4-5",
    api_base="http://127.0.0.1:8788",
    api_key="sk-ant-…",
    messages=[{"role": "user", "content": "Hello!"}],
)

Running the standalone LiteLLM Proxy instead?

Same mechanism, one YAML field: point each model's api_base at distil proxy, and every request routed through the LiteLLM Proxy is compressed transparently.

# Terminal 1 — distil sits in front of the real upstream
distil proxy --port 8788 --upstream https://api.anthropic.com

# config.yaml — Terminal 2
model_list:
  - model_name: claude-opus-4-5
    litellm_params:
      model: anthropic/claude-opus-4-5
      api_base: http://127.0.0.1:8788
      api_key: os.environ/ANTHROPIC_API_KEY

# Terminal 2 — start the LiteLLM Proxy against that config
litellm --config config.yaml

Cursor

Settings → enable Override OpenAI Base URL, set it to the proxy, and add your model. The key field is labelled "OpenAI API Key" but is sent to whatever endpoint you configure.

Partial coverage — know this before you rely on it. Cursor routes its agent and chat panel through the overridden base URL, but tab completion and ⌘K inline edit stay on Cursor's own backend and never reach the proxy. So your savings cover the long-context agent traffic (which is where the tokens are) and not the keystroke-level features. Also note that turning this on has been reported to break Anthropic BYOK models with 422s, since Claude traffic is then sent OpenAI-shaped to your endpoint.

CrewAI

CrewAI's LLM takes base_url directly. Point it at the proxy and prefix the model with its provider.

$ pip install crewai
$ distil proxy --port 8788 --upstream https://api.openai.com &

import os
from crewai import Agent, LLM

llm = LLM(
    model="openai/gpt-4o",               # provider prefix is required
    base_url="http://127.0.0.1:8788/v1",  # ← the only change; end at /v1
    api_key=os.environ["OPENAI_API_KEY"],
)
agent = Agent(llm=llm, ...)
Pass the LLM everywhere, not just to the agents. CrewAI resolves planning_llm, function_calling_llm and any manager LLM separately — leave those unset and those calls go straight to the provider, bypassing the proxy entirely. You would see partial savings and no error. Set base_url to the /v1 root, not /v1/chat/completions; LiteLLM appends the route itself.

Agno

Agno's OpenAILike is the class it recommends for any OpenAI-compatible endpoint — it takes the same parameters as OpenAIChat and relaxes the OpenAI-only assumptions. Point base_url at the proxy and every model call the agent makes is compressed on the way out.

$ pip install agno openai
$ distil proxy --port 8788 --upstream https://api.openai.com &

import os
from agno.agent import Agent
from agno.models.openai.like import OpenAILike

agent = Agent(
    model=OpenAILike(
        id="gpt-4o",
        api_key=os.environ["OPENAI_API_KEY"],
        base_url="http://127.0.0.1:8788/v1",   # ← the only change
    )
)
agent.print_response("...")

Strands Agents

Strands' OpenAIModel forwards client_args straight to the underlying openai client, so base_url goes there. Nothing else in the agent changes.

$ pip install strands-agents
$ distil proxy --port 8788 --upstream https://api.openai.com &

import os
from strands import Agent
from strands.models.openai import OpenAIModel

model = OpenAIModel(
    client_args={
        "api_key": os.environ["OPENAI_API_KEY"],
        "base_url": "http://127.0.0.1:8788/v1",  # ← the only change
    },
    model_id="gpt-4o",
)
agent = Agent(model=model)

Microsoft AutoGen

AutoGen's OpenAIChatCompletionClient takes base_url directly, same as the raw OpenAI client it wraps.

$ pip install "autogen-agentchat" "autogen-ext[openai]"
$ distil proxy --port 8788 --upstream https://api.openai.com &

import os
from autogen_agentchat.agents import AssistantAgent
from autogen_ext.models.openai import OpenAIChatCompletionClient

client = OpenAIChatCompletionClient(
    model="gpt-4o",
    api_key=os.environ["OPENAI_API_KEY"],
    base_url="http://127.0.0.1:8788/v1",  # ← the only change
)
agent = AssistantAgent("assistant", model_client=client)

Prefer no sidecar at all? distil.integrations.autogen wraps a ChatCompletionClient or a FunctionTool callable in-process — see AutoGen × Distil for both.

Why there is no distil-agno, distil-strands, or distil-autogen package. All three are two-line base_url redirects, and so is every other framework that speaks an OpenAI- or Anthropic-shaped API. distil is a proxy, so a framework is supported the moment it lets you set an endpoint — there is nothing to version, nothing to break on the framework's next release, and no per-framework code to maintain. The in-process hooks further down (and AutoGen's own page) exist only for the cases where you want compression without running a proxy at all.

LlamaIndex

LlamaIndex's OpenAI/Anthropic LLM classes take api_base directly, same as the raw client each one wraps.

$ pip install llama-index
$ distil proxy --port 8788 --upstream https://api.openai.com &

from llama_index.llms.openai import OpenAI

llm = OpenAI(model="gpt-5", api_base="http://127.0.0.1:8788/v1")  # ← the only change
query_engine = index.as_query_engine(llm=llm)

Prefer no sidecar at all? distil.integrations.llamaindex gives you a node postprocessor for retrieved context, an LLM wrapper, and a FunctionTool callable wrapper, all in-process — see LlamaIndex × Distil for all three.

Vercel AI SDK (TypeScript)

import { createAnthropic } from "@ai-sdk/anthropic";
import { generateText } from "ai";

const anthropic = createAnthropic({
  baseURL: "http://127.0.0.1:8788",
  apiKey: process.env.ANTHROPIC_API_KEY,
});

const { text } = await generateText({
  model: anthropic("claude-opus-4-5"),
  prompt: "Hello!",
});

LangChain.js (TypeScript)

import { ChatAnthropic } from "@langchain/anthropic";

const model = new ChatAnthropic({
  model: "claude-opus-4-5",
  apiKey: process.env.ANTHROPIC_API_KEY,
  anthropicApiUrl: "http://127.0.0.1:8788",
  // older versions: clientOptions: { baseURL: "http://127.0.0.1:8788" }
});

const response = await model.invoke([
  ["human", "Hello!"],
]);

Google Gemini REST

# Start proxy pointing at the Gemini API
distil proxy --port 8788 --upstream https://generativelanguage.googleapis.com
import google.generativeai as genai

genai.configure(
    transport="rest",
    client_options={"api_endpoint": "http://127.0.0.1:8788"},
)
model = genai.GenerativeModel("gemini-1.5-pro")
response = model.generate_content("Hello!")

See examples/python_gemini.py for a full runnable example including tool-use turns. The proxy transparently compresses text parts (Tier-0 lossless) and functionResponse payloads (Tier-1 reversible digest). functionCall, inlineData, fileData, model-authored text, and systemInstruction are always passed through unchanged.


Claude Code plugin

The first-class integration: a session-first savings status line plus slash commands, shipped as a marketplace-format plugin (plugins/distil). Wire it in one step with distil setup (or distil onboard), then route the agent through compression with distil wrap -- claude.

CommandWhat it does
/distil-onboardSet up distil + a guided, tailored tour
/distilSavings report + how to route more traffic through distil
/distil-statsFull breakdown — tokens, cost, runs, per-trajectory bars
/distil-shadowLive decision-equivalence: did compression preserve the next action?
/distil-dashboardHTML savings page — session, lifetime, and decision-equivalence cards
/distil-doctorDiagnose the setup — ledger, shadow validation, proxy round-trip, wiring
/distil-certifyTrajectory-level certificate: bound how many solvable tasks compression may cost
/distil-badgeShareable badge of your measured savings

One pattern in every state: distil · <live> · total ▼<lifetime>. The live segment aggregates recent activity (last 15 min, all terminals), so it never flickers between sessions:

StateYou seeMeans
savingdistil · ▼12.0K · 40% smaller · $0.31 · total ▼27.0M · ⚠de 97.5% (398)compressing your recent traffic
watchingdistil · ✓ on · waiting for a large read · total ▼27.0Mon, but no large content yet — savings come from big file/command output
idledistil · ✓ on · total ▼27.0Mset up and on, no recent traffic
not routeddistil · off — session not routed · total ▼27.0Mthis session goes straight to the provider — start it with distil wrap (or the always-on env) to compress. "on" always means routed, never merely installed.

▼ = tokens saved · total = lifetime · de = decision-equivalence (shown only past the reporting floor, 50 A/B + 30 A/A shadow samples — a rate over a handful is noise; ✓ at 99% and above, ⚠ under 99%, ✗ under 95%). Sharing the line with git/cwd/model? Set DISTIL_STATUSLINE=minimal for a two-fact segment: distil ▼75.0K · 27.0M total.


MCP server

A zero-dependency Model Context Protocol server (stdlib only — no SDK) exposes distil's reversible compression to any MCP client (Claude Desktop, IDEs, agents) over stdio:

claude mcp add distil -- distil mcp   # Claude Code — one line, done

Every other client takes the same stdio config — Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json), Cursor (.cursor/mcp.json), VS Code (.vscode/mcp.json):

{
  "mcpServers": {
    "distil": { "command": "distil", "args": ["mcp"] }
  }
}

No distil install? Run it from PyPI on demand — "command": "uvx", "args": ["--from", "distil-llm", "distil", "mcp"]. Restart the client after editing. To check the server without a client at all:

echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | distil mcp
ToolWhat it doesCalled when
distil_compress(text)Compact digest + an 8-hex handle; the original is stored locally (encrypted, 0600) and never returned until asked A tool returned something huge and carrying it verbatim is wasteful
distil_expand(handle)The exact original bytes — not a summaryThe digest lost a detail the agent now needs: a line, a value, a stack frame
distil_savings()Cumulative tokens/dollars from the local ledger Someone asks what distil has saved

Each tool carries MCP annotations (readOnlyHint, idempotentHint, openWorldHint: false), so a client knows distil_expand is a safe, repeatable, offline read instead of inferring it from prose.

This is the recall path, not the savings path. The MCP server does not compress your agent's traffic — distil wrap does that transparently, with no tool calls. What the MCP server adds is the other half: any agent, including one you never wrapped, can call distil_expand on a handle it sees in context and get the original back. Handles survive restarts and cross processes, and age out after DISTIL_RESTORE_TTL_DAYS (default 14).


In-process hooks (LiteLLM · LangChain · LangGraph · AutoGen · LlamaIndex)

Prefer not to run a sidecar? Compress the request in-process — the same reversible compression, no proxy. Every helper lazy-imports (or duck-types) its framework, so distil stays zero-dep.

# LiteLLM — drop-in for litellm.completion
from distil.integrations import litellm as distil_litellm
resp = distil_litellm.completion(model="claude-opus-4-8", messages=[...],
                                 distil_verbatim=True)  # optional, Tier-0 only

# LangChain — compress a message list before the model call (duck-typed)
from distil.integrations.langchain import compress_messages
msgs = compress_messages(state["messages"], verbatim=True)

# LangGraph — compress graph state right before the model node
from distil.integrations.langgraph import pre_model_hook
agent = create_react_agent(model, tools, pre_model_hook=pre_model_hook())

# AutoGen — compress a ChatCompletionClient's outgoing messages, or one tool's output
from distil.integrations.autogen import DistilModelClient, compressing_tool
client = DistilModelClient(OpenAIChatCompletionClient(model="gpt-4o"))
tool = FunctionTool(compressing_tool(get_weather), description="...")

# LlamaIndex — compress retrieved nodes, an LLM's outgoing calls, or one tool's output
from distil.integrations.llamaindex import DistilNodePostprocessor, DistilLLM
query_engine = index.as_query_engine(
    llm=DistilLLM(OpenAI(model="gpt-5")),
    node_postprocessors=[DistilNodePostprocessor()],
)

Tool/function messages and retrieved nodes get the reversible Tier-1 digest; human/system messages get Tier-0 lossless; the model's own words are never rewritten. compress() (LiteLLM), compress_messages() (LangChain), pre_model_hook() (LangGraph), DistilModelClient/compressing_tool() (AutoGen), and DistilLLM/DistilNodePostprocessor (LlamaIndex) are framework-free and unit-tested. The LangGraph hook returns only the updated message list, so every other state field is left intact; DistilModelClient delegates every attribute it doesn't wrap straight to the real client, and DistilLLM re-types the LLM you pass as a transparent subclass of its own class, so it still satisfies the isinstance checks LlamaIndex performs on an llm= argument.

LangChain / LangGraph as a package — langchain-distil

The two hooks above are also published as a standalone package, listed in LangChain’s own community middleware integrations. It is a thin wrapper over exactly those hooks — same certified compression path, nothing re-implemented — and it pulls distil-llm in as a dependency.

$ pip install langchain-distil

from langchain_distil import compress_messages, pre_model_hook, as_runnable

msgs  = compress_messages(msgs)                                   # a message list
graph = create_react_agent(model, tools, pre_model_hook=pre_model_hook())  # LangGraph state
chain = as_runnable() | llm                                       # or a chain step

as_runnable() imports langchain-core lazily, so importing the package never requires it. See LangChain × Distil or LangGraph × Distil for the full page.


ASGI middleware — when you host the endpoint

Every hook above sits in a client that calls out to a provider. If instead your own backend is the thing building the provider request — a FastAPI/Starlette/Litestar app that forwards to Anthropic/OpenAI/Gemini itself — wrap it once with DistilMiddleware, pure ASGI, no Starlette import:

from distil.integrations.asgi import DistilMiddleware

app = DistilMiddleware(app)  # wraps any ASGI 3 app; compresses matching POST bodies

It reuses the exact path detection and reversible compression distil proxy uses, so a handle minted here expands anywhere. See ASGI Middleware × Distil for the full page.


Observability headers

The proxy adds up to 9 response headers per compressed response. The first two appear on every compressed request; the rest are conditional:

HeaderMeaningCondition
x-distil-compressed: 1 Compression was applied this turn. Always (on compressed requests)
x-distil-tokens-saved: <n> Estimated input tokens saved (heuristic tokenizer). Always (on compressed requests)
x-distil-expanded: 1 A digest was resolved via the transparent expand loop. When --expand fired
x-distil-cache-prefix-msgs: <n> Leading messages left byte-identical vs the previous turn (the prompt-cache-read region) — the verifiable benefit of a prefix-freeze router, content-free. With --session-delta
x-distil-cache-refs: <n> Total cache-delta references this turn. With --session-delta
x-distil-cache-delta: <n> Delta-encoded references this turn. With --session-delta
x-distil-cache-tokens-saved: <n> Tokens saved by cache-delta encoding. With --session-delta
x-distil-output-shaping: light|aggressive Output shaping was injected at this level. When --shape-output fired
x-distil-shadow: sampled This request was sampled for shadow-mode decision-equivalence. When --shadow rate triggered

The managed gateway (distil gateway) additionally adds x-distil-tenant: <id> for per-tenant accounting.


Node.js / TypeScript

Distil's universal path is the proxy — it's language-agnostic, so JS/TS stacks need no Distil-specific package. Start the proxy (or wrap your agent) and point your SDK's baseURL at it:

# start the proxy (needs Python 3.9+ — see Install)
distil proxy --port 8788 --upstream https://api.anthropic.com
// then in your Node/TS app — no code change beyond baseURL
const client = new Anthropic({ baseURL: "http://localhost:8788" });

Or wrap an existing CLI/agent in one shot: distil wrap -- <your-command> starts the proxy and injects ANTHROPIC_BASE_URL automatically.

Homebrew

brew tap dshakes/tap
brew install dshakes/tap/distil
distil proxy --port 8788

Adapters & Integration

Three paths to production: a one-line Python wrapper that compresses in-process, a standalone HTTP proxy that works with any SDK or framework, or distil wrap — a zero-config launcher that routes any existing command through the proxy automatically.


Path 1 — In-process: wrap(client)

Module: distil/adapters/anthropic.py

The fastest path for Python codebases already using the Anthropic SDK. Wrap your client once at construction time — all downstream messages.create calls are transparently compressed and cache-pinned with no call-site changes.

import anthropic
from distil.adapters.anthropic import wrap

# Before: plain Anthropic client
client = anthropic.Anthropic()

# After: one-line drop-in
client = wrap(anthropic.Anthropic())

# No other code changes — all calls compressed transparently
response = client.messages.create(
    model="claude-opus-4-5",
    max_tokens=1024,
    system="You are a helpful assistant.",
    messages=[{"role": "user", "content": "Analyse this log output…"}],
)

What the wrapper does

On every messages.create(**kwargs) call, the wrapper:

  1. Compresses the messages array via compress_messages():
    • User text blocks: Tier-0 lossless transforms (JSON minify + run collapse).
    • Tool result blocks with ≥ 6 lines: Tier-1 reversible digest, original stored in a local RestoreStore.
    • Assistant text and tool_use blocks: passed through unchanged.
    • image blocks: a duplicate of an already-seen image (certificate-gated, ADR 0003) is elided to a reversible reference; the first occurrence is always left untouched.
  2. Pins the cache prefix via place_cache_control(): marks the last stable system block with {"cache_control": {"type": "ephemeral"}}, so the first call pays the cache-write price and every subsequent call pays the ~0.1× cache-read price.
  3. Forwards the modified kwargs to the real client.messages.create.

The RestoreStore

Digested blocks embed an 8-hex SHA-256 handle in their compressed text. The RestoreStore maps handles to originals locally — it is never sent to the model and costs zero tokens. To recover an original:

from distil.adapters.anthropic import compress_messages

compressed, store = compress_messages(messages)
# store.handles → frozenset of all active handles
original = store.expand("a3f92b1c")  # retrieve original by handle

Design properties


Path 2 — Provider proxy: distil proxy

Module: distil/proxy.py

A lightweight HTTP proxy that sits between your client and the real LLM API. It intercepts POST /v1/messages, POST /v1/chat/completions, POST /v1/responses, and POST /v1beta/models/{model}:generateContent (and :streamGenerateContent) requests, compresses the payload, and forwards the modified request to the upstream. All other paths and methods pass through unchanged.

This approach works with any SDK, framework, or language, with zero application code changes.

$ distil proxy --port 8788 --upstream https://api.anthropic.com
distil proxy listening on http://127.0.0.1:8788
  → upstream: https://api.anthropic.com

Point your client at the proxy

Anthropic SDK

Python

import anthropic
client = anthropic.Anthropic(
    base_url="http://localhost:8788"
)
OpenAI SDK

Python / Node

import openai
client = openai.OpenAI(
    base_url="http://localhost:8788/v1",
    api_key="…"
)
LiteLLM

Multi-provider

litellm.completion(
    api_base="http://localhost:8788",
    model="anthropic/claude-…",
    messages=[...]
)
LangChain

Python

from langchain_anthropic import ChatAnthropic
llm = ChatAnthropic(
    base_url="http://localhost:8788"
)
Google Gemini REST

Python / curl

$ distil proxy --upstream \
    https://generativelanguage.googleapis.com

# then in your code — no other changes
import google.generativeai as genai
genai.configure(
    transport="rest",
    client_options={
        "api_endpoint": "http://localhost:8788"
    },
)

Proxy flags

FlagDefaultDescription
--port8788Port to listen on
--upstreamhttps://api.anthropic.comReal LLM API base URL (no trailing slash)
--lossless-onlyoffPolicy mode: Tier-0 verbatim only — no lossy output-shaping, no tool injection (--expand), no Tier-1 digest stubs. Folds directly into verbatim (no separate --verbatim needed). Use for subscription / OAuth sessions.
--verbatimoffSkip the Tier-1 digest entirely; Tier-0 only (JSON minify + collapse exact-duplicate runs). The model sees message content essentially verbatim. Use for interactive sessions or where distil_expand is unavailable. Lower savings.

Response headers

The proxy adds up to 9 response headers to compressed responses (not all appear on every request):

The managed gateway additionally adds x-distil-tenant: <id> for per-tenant accounting.

Threading model

The proxy uses Python's ThreadingHTTPServer — each incoming connection is handled in its own thread, so concurrent requests from multiple clients are supported. It is intentionally simple and easy to audit.


Path 3 — Transparent: distil wrap

Module: distil/proxy.py · wrap_run()

When you don't own the agent's code — it's a CLI, a third-party tool, or any process that reads ANTHROPIC_BASE_URL — distil wrap gives you Path 2's compression with none of the setup. It spawns the proxy on an ephemeral port, injects the env var into the child, runs your command, and tears everything down on exit (flushing genuine savings to your ledger).

$ distil wrap -- claude -p "summarize this repo"
distil wrap → proxy http://127.0.0.1:54xxx (upstream https://api.anthropic.com)
  → ANTHROPIC_BASE_URL=http://127.0.0.1:54xxx
  → recording genuine savings → distil leaderboard

Everything after -- runs verbatim and its exit code is propagated. Use --env-var to point a different variable (e.g. an OpenAI-compatible client) at the proxy. See the CLI reference for all flags.


Google Gemini adapter

Module: distil/adapters/gemini.py

The Gemini adapter compresses Google's generateContent REST request shape. It is wired into the proxy, the async proxy, and the gateway — no extra configuration beyond pointing the proxy at the Gemini API.

$ distil proxy --upstream https://generativelanguage.googleapis.com
distil proxy listening on http://127.0.0.1:8788
  → upstream: https://generativelanguage.googleapis.com

See examples/python_gemini.py for a runnable end-to-end example.

Request shape handled

{
  "contents": [
    {
      "role": "user" | "model",
      "parts": [
        { "text": "…" },
        { "functionCall": {"name": "…", "args": {…}} },
        { "functionResponse": {"name": "…", "response": {…}} }
      ]
    }
  ]
}

What is compressed

Part typeWhat happensReversible?
text parts (non-model role) Tier-0 lossless: JSON minify + collapse exact-duplicate runs Provably lossless — no side state
functionResponse parts Large string values inside response get the Tier-1 reversible digest; object structure is preserved so the request stays valid Reversible via local RestoreStore (never sent to model, zero tokens)
functionCall parts Passed through untouched —
text parts (model role) Passed through untouched — the model's own words are never rewritten —
inlineData, fileData (non-model role) A duplicate of an already-seen image is elided to a reversible reference (certificate-gated, ADR 0003, same rule as the Anthropic/OpenAI adapters); first occurrence untouched. A fileData.fileUri is only treated as a duplicate when it embeds a data: URI — a plain URL is never proof of identical bytes. Reversible via local RestoreStore (never sent to model, zero tokens)
systemInstruction Left byte-exact (same treatment as the Anthropic system field) —

Path detection

The adapter activates on:

Savings reporting

Token savings are reported via the x-distil-tokens-saved response header and recorded to the savings ledger, exactly as for other adapters.

Shadow-mode decision-equivalence

Shadow-mode live decision-equivalence works for Gemini. It reads the chosen action from candidates[0].content.parts[].functionCall in the response.

Output verbosity shaping

Pass --shape-output to the proxy and it will inject a conciseness directive into systemInstruction using the Gemini-specific shaping path (shape="gemini" internally). The same --lossless-only guard that blocks shaping on subscription sessions applies here too.

Documented seams (not yet wired)

Expand-tool injection (--expand) is not yet wired for the Gemini tool shape — the scaffolding exists but the distil_expand tool is not injected into Gemini requests. Gemini context caching (cachedContent) is also not yet wired. These gaps are documented honestly here rather than papered over.

Verbatim and lossless-only modes

--verbatim is the Tier-0-only mode for all adapters, including Gemini: the Tier-1 digest is skipped entirely, so the model sees message content byte-for-byte (JSON minify + exact-duplicate-run collapse only). Use it for interactive (human-in-the-loop) sessions, out-of-distribution traffic, or anywhere distil_expand recovery is unavailable. Lower savings.

--lossless-only is the policy mode: it restricts the proxy to Tier-0 verbatim — no lossy output-shaping, no tool injection (--expand), and no Tier-1 digest stubs either (without an expand tool the agent cannot recover a stub, so the flag folds directly into verbatim). This is the correct flag for subscription or OAuth-gated deployments; no separate --verbatim is needed.


Choosing your path

wrap(client)distil proxy
LanguagePython onlyAny (HTTP)
SDK requiredAnthropic SDK (duck-typed)Any base_url-honoring client
SetupOne import + one callStart proxy process, update base_url
RestoreStore accessDirect Python objectNot exposed (stateless proxy)
Cache-control pinningAutomatic via place_cache_control()Not currently applied
lossless-only (policy) modeNot a compress_messages parameter — enforce at the call site by omitting the call when policy requires it--lossless-only flag
verbatim (Tier-0 only) modePass verbatim=True to compress_messages--verbatim flag
Multi-language teamsNoYes

Honest note: live provider path

The proxy forwards to the real upstream. Any request that goes through distil proxy will hit the live provider API and consume your quota / billing. Ensure you have a valid API key set in the client — the proxy does not add or modify authorization headers.

The --runner anthropic flag on distil certify also routes to the live Anthropic API. It is implemented and wired up, but marked UNVERIFIED in the README until you run it with a real key. The offline deterministic runner (the default) is what powers all corpus measurements.


Library API

Embed distil in your own agent — Python or TypeScript, in-process, no proxy and no daemon.

The proxy is still the zero-config path, and it is the one that reaches the reversible digest tier. Reach for the library when you are building the agent: you already hold the message list, and you want compression without a process to keep alive — a serverless function, an edge worker, a CI job.

Python

from distil import compress_messages, expand_handle

result = compress_messages(messages)          # OpenAI/Anthropic-style dicts
print(f"{result.saved_pct:.1f}% smaller")
response = client.messages.create(model=..., messages=result.messages)

original = expand_handle(result.handles[0])   # byte-exact, any time, any process

Tool results get the reversible digest; user and system text get lossless transforms only; the model's own turns are never rewritten. The input list is never mutated, and unchanged messages come back by identity. Pass verbatim=True for lossless-only with no handles.

Named compress_messages/expand_handle rather than compress/expand because distil.compress and distil.expand are modules. Python binds a submodule onto its parent on import, so a top-level export with those names would resolve to the function in a fresh interpreter and to the module in any program that had touched the submodule. A test enumerates submodules and fails on any future collision.

TypeScript

import { compress } from "distil-llm";

const r = compress(messages);                 // never mutates the input
console.log(`${r.savedPct.toFixed(1)}% smaller`);

And as Vercel AI SDK middleware:

import { wrapLanguageModel } from "ai";
import { distilMiddleware } from "distil-llm";

const model = wrapLanguageModel({ model: myModel, middleware: distilMiddleware() });

The middleware implements transformParams only. Compression happens on the way in; wrapping the response would mean rewriting the model's own output. Tool results are handled in both the v5 (output.value) and v4 (result) shapes, so an SDK bump cannot silently stop compressing the largest thing in an agent's context.

The tier boundary — and why it is where it is

PathTierCertificate
distil wrapdigest (reversible)✅ covered
proxy / gatewaydigest (reversible)✅ covered
MCP serverdigest (reversible)✅ covered
in-process librarylossless onlyn/a — nothing is elided

The in-process libraries are lossless-tier on purpose. The digest mints restore handles whose originals must live in one store shared with the proxy and the MCP server — otherwise expand fails to resolve a handle the model can see. It is also the tier the decision-equivalence certificate measures, so a second implementation of it would be a second thing to certify, and nobody would notice when the two drifted.

The TypeScript port is held byte-identical

A JS implementation that merely compressed similarly would hand you a guarantee that does not describe the code you are running. So the port is checked against the Python engine over a shared corpus, and CI fails on any divergence.

Where byte-identity is not achievable the port declines rather than emitting output the certificate does not cover — Python renders an integral float as 1.0 and JS renders it 1; JS objects hoist integer-like keys. Both are detected on the source and the text is left byte-exact. Tests assert the declines and assert that safe cases still compress, so conformance cannot be achieved by declining everything.

Framework hooks

FrameworkEntry point
LangChaindistil.integrations.langchain.compress_messages
LangGraphdistil.integrations.langgraph.pre_model_hook()
LiteLLMdistil.integrations.litellm.compress(kwargs)
Agnodistil.integrations.agno.compressed_model(model)
Strandsdistil.integrations.strands.compressing_hook()
Vercel AI SDKdistilMiddleware() (npm)

Every one is duck-typed and never imports its framework. That is what keeps distil a zero-dependency install, and it means a framework release cannot break the integration.

Runnable examples

Each of these runs as-is: python_library.py · js_library.ts · js_ai_sdk_middleware.ts · python_agno.py · python_strands.py


Anthropic SDK × Distil

One line: point the client's base_url at the proxy. Streaming, tool use and prompt caching all pass through unchanged.

Setup

Start the proxy against the provider this SDK talks to, then point the client at it. Nothing else in your code changes.

$ distil proxy --port 8788 --upstream https://api.anthropic.com &

import anthropic

client = anthropic.Anthropic(
    base_url="http://127.0.0.1:8788",   # ← the only change
)

There is also an in-process path that needs no proxy at all: distil.adapters.anthropic.wrap(client) compresses messages.create calls directly. Use it when you cannot run a sidecar.

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Distil is a proxy, so anything that lets you set a base URL works — this page is a worked example, not the boundary of support.


OpenAI SDK × Distil

Point base_url at the proxy's /v1 root. Chat Completions and the Responses API are both intercepted.

Setup

Start the proxy against the provider this SDK talks to, then point the client at it. Nothing else in your code changes.

$ distil proxy --port 8788 --upstream https://api.openai.com &

import os
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:8788/v1",  # ← the only change
    api_key=os.environ["OPENAI_API_KEY"],
)
Worth knowing. End the base URL at /v1, not /v1/chat/completions — the SDK appends the route itself.

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Distil is a proxy, so anything that lets you set a base URL works — this page is a worked example, not the boundary of support.


LiteLLM × Distil

LiteLLM takes api_base. The same proxy also works behind the standalone LiteLLM Proxy — one YAML field per model.

Setup

Start the proxy against the provider this SDK talks to, then point the client at it. Nothing else in your code changes.

$ distil proxy --port 8788 --upstream https://api.openai.com &

import litellm

resp = litellm.completion(
    model="gpt-4o",
    api_base="http://127.0.0.1:8788/v1",   # ← the only change
    messages=[{"role": "user", "content": "..."}],
)

Running the standalone LiteLLM Proxy instead? Set each model's api_base to the distil proxy in your config YAML; every request routed through LiteLLM is then compressed transparently.

In-process: the LiteLLM Proxy hook

No sidecar needed on the standalone LiteLLM Proxy. Install the extra and register distil's async_pre_call_hook, which compresses each request just before LiteLLM routes it:

$ pip install 'distil-llm[litellm]'

# config.yaml
litellm_settings:
  callbacks: distil.integrations.litellm_hook.proxy_handler_instance

Or choose options in your own callback file: proxy_handler_instance = DistilCompressionHook(digest=False).

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Distil is a proxy, so anything that lets you set a base URL works — this page is a worked example, not the boundary of support.


LangChain × Distil

Two routes. The middleware package compresses the message list in your own process — no proxy, no network hop. The proxy needs no code change at all beyond a base URL. Pick one; they are not meant to be stacked.

Middleware — langchain-distil

Listed in LangChain’s own community middleware integrations. If you arrived from there, this is the package.

$ pip install langchain-distil

from langchain_distil import compress_messages, pre_model_hook, as_runnable

# 1. compress a message list directly
msgs = compress_messages(msgs)

# 2. LangGraph: compress graph state before the model node
graph = create_react_agent(model, tools, pre_model_hook=pre_model_hook())

# 3. or drop it into a chain (langchain-core is imported lazily)
chain = as_runnable() | llm

Tool and function messages get the reversible Tier-1 digest, human and system messages are Tier-0 lossless, and assistant messages are never rewritten — a model’s own words are not distil’s to edit. Every digest is byte-exact recoverable through distil_expand. Pass verbatim=True for Tier-0 only, which is the right choice when no recovery tool is wired up.

It is a thin wrapper over distil.integrations.langchain and distil.integrations.langgraph, so it inherits the same certified compression path rather than re-implementing it. distil-llm comes along as a dependency — you do not install both by hand.

Setup — the proxy route

Start the proxy against the provider this SDK talks to, then point the client at it. Nothing else in your code changes.

$ distil proxy --port 8788 --upstream https://api.anthropic.com &

from langchain_anthropic import ChatAnthropic

llm = ChatAnthropic(
    model="claude-opus-4-8",
    base_url="http://127.0.0.1:8788",     # ← the only change
)

LangGraph users can compress graph state directly instead, with distil.integrations.langgraph.pre_model_hook() — no proxy required.

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Distil is a proxy, so anything that lets you set a base URL works — this page is a worked example, not the boundary of support.


LangGraph × Distil

Two routes. pre_model_hook compresses graph state in your own process — no proxy, no network hop. The proxy needs no code change at all beyond a base URL. Pick one; they are not meant to be stacked.

In-process — distil.integrations.langgraph

Duck-typed: this module never imports langgraph or langchain, so it costs nothing to import if you end up not using it. pre_model_hook() plugs straight into LangGraph’s own seam for transforming state right before the model node.

from distil.integrations.langgraph import pre_model_hook

graph = create_react_agent(model, tools, pre_model_hook=pre_model_hook())

# or, manually, inside any node:
from distil.integrations.langgraph import compress_state
state = compress_state(state, verbatim=True)

Both helpers work on a dict-like state (state["messages"]) or an attribute-style state (state.messages); a state with no message list is returned untouched. pre_model_hook returns only {"messages": ...}, so every other field on the graph state is left exactly as LangGraph produced it.

Tool and function messages get the reversible Tier-1 digest, human and system messages are Tier-0 lossless, and assistant messages are never rewritten — a model’s own words are not distil’s to edit. Every digest is byte-exact recoverable through distil_expand. Pass verbatim=True for Tier-0 only, which is the right choice when no recovery tool is wired up.

The same two hooks are also published as the standalone langchain-distil package, listed in LangChain’s own community middleware integrations — see LangChain × Distil if you arrived from there.

Setup — the proxy route

Start the proxy against the provider this SDK talks to, then point the client at it. Nothing else in your code changes — including graphs that call the model directly rather than through create_react_agent.

$ distil proxy --port 8788 --upstream https://api.anthropic.com &

from langchain_anthropic import ChatAnthropic

llm = ChatAnthropic(
    model="claude-opus-4-8",
    base_url="http://127.0.0.1:8788",     # ← the only change
)

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone. The in-process hook has no proxy to check against — distil stats after a run is the equivalent signal.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Distil is a proxy, so anything that lets you set a base URL works — this page is a worked example, not the boundary of support.


Vercel AI SDK × Distil

Two routes. distilMiddleware() compresses the request in your own process — no proxy, no network hop. The proxy needs no code change at all beyond a base URL — and streaming works unchanged either way, since the proxy relays the stream rather than buffering it. Pick one; they are not meant to be stacked.

Middleware — distilMiddleware()

Wraps wrapLanguageModel's own middleware seam, from the distil-llm npm package.

$ npm install distil-llm

import { wrapLanguageModel } from "ai";
import { distilMiddleware } from "distil-llm";

const model = wrapLanguageModel({
  model: gateway("anthropic/claude-sonnet-5"),
  middleware: distilMiddleware({
    onSavings: (s) => console.log(`${s.savedPct.toFixed(1)}% smaller`),
  }),
});

Implements transformParams only — compression happens on the way in, and wrapping the response would mean rewriting the model’s own output, which distil does not do. Tool results (both the v4 result and v5 output.value shapes), user text parts, and system strings are compressed; assistant messages, files, and tool calls are passed through untouched.

Lossless tier only, like the package’s compress() helper. Route through the proxy below for the reversible digest tier.

Setup — the proxy route

Start the proxy against the provider this SDK talks to, then point the client at it. Nothing else in your code changes.

$ distil proxy --port 8788 --upstream https://api.anthropic.com &

import { createAnthropic } from "@ai-sdk/anthropic";

const anthropic = createAnthropic({
  baseURL: "http://127.0.0.1:8788",        // ← the only change
});

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Distil is a proxy, so anything that lets you set a base URL works — this page is a worked example, not the boundary of support.


Agno × Distil

Agno's OpenAILike is the class it recommends for any OpenAI-compatible endpoint — same parameters as OpenAIChat, without the OpenAI-only assumptions.

Setup

Start the proxy against the provider this SDK talks to, then point the client at it. Nothing else in your code changes.

$ distil proxy --port 8788 --upstream https://api.openai.com &

import os
from agno.agent import Agent
from agno.models.openai.like import OpenAILike

agent = Agent(
    model=OpenAILike(
        id="gpt-4o",
        api_key=os.environ["OPENAI_API_KEY"],
        base_url="http://127.0.0.1:8788/v1",  # ← the only change
    )
)

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Distil is a proxy, so anything that lets you set a base URL works — this page is a worked example, not the boundary of support.


Strands Agents × Distil

Strands forwards client_args straight to the underlying openai client, so the base URL goes there. Nothing else in the agent changes.

Setup

Start the proxy against the provider this SDK talks to, then point the client at it. Nothing else in your code changes.

$ distil proxy --port 8788 --upstream https://api.openai.com &

import os
from strands import Agent
from strands.models.openai import OpenAIModel

model = OpenAIModel(
    client_args={
        "api_key": os.environ["OPENAI_API_KEY"],
        "base_url": "http://127.0.0.1:8788/v1",  # ← the only change
    },
    model_id="gpt-4o",
)
agent = Agent(model=model)

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Distil is a proxy, so anything that lets you set a base URL works — this page is a worked example, not the boundary of support.


Microsoft AutoGen × Distil

OpenAIChatCompletionClient takes a base_url like any OpenAI-compatible client, so the proxy route needs no code beyond that. A second, in-process module also ships for teams that don't want a sidecar at all.

Setup — the proxy route

Start the proxy against the provider this SDK talks to, then point the client at it. Nothing else in your code changes.

$ pip install "autogen-agentchat" "autogen-ext[openai]"
$ distil proxy --port 8788 --upstream https://api.openai.com &

import os
from autogen_agentchat.agents import AssistantAgent
from autogen_ext.models.openai import OpenAIChatCompletionClient

client = OpenAIChatCompletionClient(
    model="gpt-4o",
    api_key=os.environ["OPENAI_API_KEY"],
    base_url="http://127.0.0.1:8788/v1",  # ← the only change
)
agent = AssistantAgent("assistant", model_client=client)

In-process — no sidecar

distil.integrations.autogen is duck-typed and never imports autogen_core, so it costs nothing to have installed either way. Two seams, matching AutoGen's own shapes (verified against the official docs, autogen-core 0.4+/0.7+):

from distil.integrations.autogen import DistilModelClient, compressing_tool

# 1. Compress a tool's return value before it becomes a FunctionExecutionResult
async def get_weather(city: str) -> str:
    return "... huge forecast ..."

tool = FunctionTool(compressing_tool(get_weather), description="Get the weather")

# 2. Or compress every outgoing model call, transparently
client = DistilModelClient(OpenAIChatCompletionClient(model="gpt-4o"))
agent = AssistantAgent("assistant", model_client=client)

DistilModelClient delegates every attribute it doesn't wrap (model_info, capabilities, count_tokens, ...) straight to the real client, so it drops in wherever a ChatCompletionClient is expected. FunctionExecutionResultMessage content gets the reversible Tier-1 digest; SystemMessage/UserMessage get Tier-0 lossless; AssistantMessage — the model's own words — is never rewritten. Pass verbatim=True to either helper for Tier-0 only.

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Hosting your own LLM-facing endpoint instead of using an SDK? See the ASGI middleware.


LlamaIndex × Distil

LlamaIndex's own OpenAI/Anthropic LLM classes take an api_base like any OpenAI/Anthropic-compatible client, so the proxy route needs no code beyond that. A second, in-process module also ships for teams that don't want a sidecar at all — including a node postprocessor for compressing retrieved context.

Setup — the proxy route

Start the proxy against the provider this SDK talks to, then point the LLM at it. Nothing else in your code changes.

$ pip install llama-index
$ distil proxy --port 8788 --upstream https://api.openai.com &

from llama_index.llms.openai import OpenAI

llm = OpenAI(model="gpt-5", api_base="http://127.0.0.1:8788/v1")  # ← the only change
query_engine = index.as_query_engine(llm=llm)

In-process — no sidecar

distil.integrations.llamaindex is duck-typed and never imports llama_index, so it costs nothing to have installed either way. Three seams, matching LlamaIndex's own shapes (verified against the official API reference and the llama-index-core source):

from distil.integrations.llamaindex import DistilNodePostprocessor, DistilLLM, compressing_tool

# 1. Compress retrieved nodes before they reach the LLM's context window
query_engine = index.as_query_engine(
    node_postprocessors=[DistilNodePostprocessor()],
)

# 2. Compress a tool's return value before an agent sees it
def get_weather(city: str) -> str:
    return "... huge forecast ..."

tool = FunctionTool.from_defaults(fn=compressing_tool(get_weather))

# 3. Or compress every outgoing chat/completion call, transparently
llm = DistilLLM(OpenAI(model="gpt-5"))
query_engine = index.as_query_engine(llm=llm)

DistilNodePostprocessor drops straight into node_postprocessors=[...]: it exposes postprocess_nodes/apostprocess_nodes, the two methods a query engine actually calls, so nothing needs to subclass BaseNodePostprocessor. Retrieved node text gets the reversible Tier-1 digest, the same as tool output elsewhere in this package. DistilLLM hands back your LLM re-typed as a transparent subclass of its own class, with only the chat/completion methods overridden and every other attribute (metadata, callback_manager, ...) untouched — so isinstance still holds, which is what resolve_llm() and every Pydantic llm: LLM field actually check; chat messages and completion prompts get Tier-0 lossless, and the model's own assistant-role replies are never rewritten. Pass verbatim=True to any of the three for Tier-0 only.

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results and retrieved nodes are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Hosting your own LLM-facing endpoint instead of using an SDK? See the ASGI middleware.


CrewAI × Distil

CrewAI's LLM takes base_url directly. Prefix the model with its provider.

Setup

Start the proxy against the provider this SDK talks to, then point the client at it. Nothing else in your code changes.

$ distil proxy --port 8788 --upstream https://api.openai.com &

import os
from crewai import Agent, LLM

llm = LLM(
    model="openai/gpt-4o",                 # provider prefix is required
    base_url="http://127.0.0.1:8788/v1",   # ← the only change; end at /v1
    api_key=os.environ["OPENAI_API_KEY"],
)
agent = Agent(llm=llm, ...)
Worth knowing. CrewAI resolves planning_llm, function_calling_llm and any manager LLM separately. Leave those unset and those calls go straight to the provider, bypassing the proxy — you would see partial savings and no error.

What actually happens

The proxy intercepts only the compressible paths — /v1/messages, /v1/chat/completions, /v1/responses and the Gemini generateContent routes. Everything else passes through untouched, and your API key travels in the request headers exactly as normal: the proxy never logs or stores it.

Large tool results are replaced by reversible digests carrying a content handle, and the original stays on your machine. The agent can pull any of it back mid-task through the distil_expand tool, so nothing is permanently discarded — which is what lets the compression be aggressive without being a gamble.

Check it before you trust it. distil simulate -m request.json replays one of your real requests through the pipeline with no model in the path and reports what would be compressed, what would be left byte-exact, and which rule protected it. See the CLI reference.

Verify it is actually routing

The most common failure is silent: the client never reaches the proxy and everything still works, just uncompressed. Two checks:

$ curl -s localhost:8788/distil/health
{"status":"ok"}

$ distil dashboard          # live savings; zero here means nothing is routing

Every response also carries x-distil-tokens-saved, so a request that went through the proxy is identifiable from its headers alone.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations. Distil is a proxy, so anything that lets you set a base URL works — this page is a worked example, not the boundary of support.


ASGI Middleware × Distil

Every other integration on this site is a client pointed at the proxy. This one is for the opposite shape: you host the LLM-facing endpoint — a FastAPI/Starlette/Litestar backend that builds a request and forwards it to Anthropic, OpenAI, or Gemini itself. Wrap your ASGI app once; its own outbound call sees a compressed body.

Setup

DistilMiddleware is pure ASGI — it never imports Starlette or FastAPI — so it wraps any ASGI 3 application, whichever framework built it.

from fastapi import FastAPI
from distil.integrations.asgi import DistilMiddleware

app = FastAPI()

@app.post("/v1/messages")
async def proxy_to_anthropic(request):
    # request.body() here already has compressed messages —
    # the middleware rewrote it before FastAPI's routing even saw it.
    ...

app = DistilMiddleware(app)                 # wrap once, at the bottom of the file
# app = DistilMiddleware(app, verbatim=True)  # Tier-0 lossless only

Under Starlette/FastAPI, wrap in main.py after the routes are declared (middleware wraps the whole ASGI callable, not a per-route decorator). Under a bare ASGI server, wrap whatever callable you pass to it: uvicorn.run(DistilMiddleware(app)).

What actually happens

The middleware inspects only POST requests whose path matches a compressible shape: /v1/messages, /v1/chat/completions, /v1/responses, or a Gemini generateContent route — the exact same detection distil proxy uses, imported rather than reimplemented. Everything else, including every other verb, passes through with the original ASGI receive untouched, at zero cost.

For a match, it drains the body, runs the same reversible compression the sidecar proxy uses (adapters.anthropic.compress_messages for the Anthropic/OpenAI shape, adapters.gemini.compress_generate_request for Gemini's), fixes up content-length, and replays the compressed bytes to your app as a normal http.request event. Digest handles land in the same on-disk restore store the proxy uses, so a handle minted here expands anywhere — distil_expand, the MCP server, or a later request through the proxy.

Fail-open by construction. A non-JSON body, an unrecognized shape, or a compression error all forward the original bytes unchanged rather than breaking the request. A body larger than the proxy's own size guard is forwarded uncompressed rather than buffered in full.

When to reach for this instead of the proxy

Use the sidecar proxy (distil proxy) whenever you can — zero code, works for any client. Reach for DistilMiddleware specifically when your own backend is the thing constructing the provider request, so there is no client base_url to redirect: the compression has to happen inside your process, on the way out.


Full matrix of every supported SDK, including the in-process hooks that need no proxy: Integrations.