Docs / Overview
Overview
Firstpass is an open-source adaptive LLM router. Every request opens on the cheapest model in a ladder; a gate checks the actual output; it serves the first output that passes and escalates one rung only on gate failure. Every decision gets a tamper-evident, hash-chained receipt.
What Firstpass is#
Prediction-based routing and its limits
Most model gateways route by prediction: a classifier reads the prompt and guesses which model can handle it, before a single token of output exists. If the guess is wrong, you find out in production — and there's no artifact explaining why the request went where it did.
How proof routing works
Firstpass routes by proof. It opens on the cheapest model, then gates the real output — runs your tests, checks a schema, asks a judge, or measures self-consistency. If the output passes, it serves. If it fails, it escalates exactly one rung and gates again. The cheap model handles most traffic; the expensive model is spent only when the cheap one is provably not enough.
The positioning is one line: proof over prediction — verified cascade routing.
The four moves#
Everything the router does reduces to four moves, in order, on every request.
What each move does
/v1/feedback; adaptive conformal moves the serve threshold under drift.How it works follows a single request through all four in detail.
POST /v1/messages (Anthropic), POST /v1/chat/completions (OpenAI), or POST /v1/responses (OpenAI Responses — including tool calls, translated both ways) your agent would send to a provider, swaps only the model field when routing, and forwards everything else verbatim. SSE streaming is buffered so the gate can run, then re-emitted as SSE — your agent sees a normal response. The four moves all happen in that buffer window, transparently.What is actually new#
Five things, and Firstpass is precise about which are novel versus assembled from known parts:
The five novelties, precisely stated
- Tamper-evident audit + independent verification. Every decision is a hash-chained receipt; an external auditor re-derives the whole chain from genesis with no access to the proxy or its database. See Receipts & audit.
- Outcome-feedback calibration of the serve threshold. Real outcomes — not prompt features — move the threshold that decides serve-versus-escalate. See Adaptive conformal.
- Zero-retrain model onboarding. Add a rung as
provider/model; there is no classifier to retrain when models change. See Providers. - Predict-to-start, verify-to-serve. A learned start-rung bandit predicts where to begin; the gate still proves what to serve. Prediction sets the entry point; proof sets the exit. See Learned start-rung.
- A served-failure guarantee. A distribution-free bound on how often a wrong answer is served, earned from a real run. See The guarantee.
Decision-model routers (Jev) vs Firstpass#
[escalation.prior] is experimental, default-off: the original sim verdict was PRIOR=STOP, a second pre-registration replayed on real MBPP data with OpenJev now reads PROCEED (see below); a third found blending a learned signal into it is BLEND-NEUTRAL; the decision gate measured against OpenJev is NOT-RECOMMENDED (hosted Jev still unmeasured).A new class of "System One" decision models — TypeSafe's Jev — answers one cheap closed-form question per query ($0.042/M input tokens, no text generation) and routes a query to a model tier by asking it to pick the tier directly. Products built on it (jev-router, prismhq/jev-router) then serve that tier's output unverified.
Firstpass takes the same cheap signal and puts it somewhere narrower: the existing start-rung bandit already had a slot for a per-query prior on where the ladder opens. The gate still runs on every attempt — a wrong prediction costs money or latency, never a wrong answer shipped.
The measured gap (simulation, not live)
Firstpass's own σ-sweep (cargo run -p firstpass-bench, n=500) puts a Jev-style unverified router's served-failure at 50.6–56.2% across the noise levels tested, against Firstpass's own gated served-failure of 15.8% on the same suite. Full numbers and both addenda: ADR 0013.
[-0.00074, -0.00031] excludes 0, served-failure held (0.0786 → 0.0778). Verdict: PROCEED. The win is ladder-dependent (−6.3% haiku→sonnet, −2.4% haiku→opus, 0% on the ~20x-price-ratio gpt-4.1-mini→gpt-5.5 ladder) and OpenJev says nothing about hosted Jev's own accuracy. [escalation.prior] stays default-off — the replay is evidence about the prior mechanism, not about the hosted vendor. Full numbers: openjev-prior-replay.md.prior+learned $0.01094 vs prior $0.01075 — paired diff +0.00019 [+0.00004, +0.00036], excludes 0, the blend is worse. Verdict: BLEND-NEUTRAL — the prior alone stays the recommendation. This run also traced an earlier "cost-aware learned-p" arm to a hindsight leak: it decided on each task's own realized cost, which only exists after generation, so its savings claim (including any earlier "~22% cheaper than first-pass" figure) is withdrawn — see ADR 0013. Full numbers: prior-blend-replay.md.[escalation.prior] — enabling it anyway
# experimental, default-off — sim was PRIOR=STOP, real-data replay (OpenJev) says PROCEED
[escalation.prior]
provider = "typesafe" # the only accepted value today
base_url = "https://api.typesafe.ai" # default; self-hosted/mock Jev overrides this
model = "jev-latest"
api_key_env = "TYPESAFE_API_KEY" # env var name only — the key itself is never logged
timeout_ms = 150 # any error fails open — no prior, no decision_prior field
strength = 10 # Beta pseudo-count weight of the prior
rungs = [
"rung 0 (claude-haiku) is the least capable tier that fully handles this request",
"rung 1 (claude-sonnet) is the least capable tier that fully handles this request",
]rungs must have exactly one entry per rung of every enforce-mode route's ladder — a length mismatch is rejected at config parse, never silently truncated. The gate still verifies every served output; the prior only ever moves where the ladder starts.decision gate — NOT-RECOMMENDED (OpenJev)
The same Jev model can also sit behind a [[gate]] block as a cheap external verifier (~$0.042/M input tokens) instead of a frontier LLM judge. A transport error, timeout, or unparseable reply ABSTAINs — never a fabricated pass. Measured on 974 served MBPP answers (111 oracle-wrong) with local OpenJev as the verifier: catch rate 0.2162 [0.1441, 0.2973], collateral 0.1031 [0.0834, 0.1228], AUC 0.6310 [0.5711, 0.6886] — below the pre-registered bar (catch ≥0.30 and collateral ≤0.05). Verdict: NOT-RECOMMENDED at τ=0.5. This measures OpenJev, not hosted Jev, which remains unmeasured. It also caught the gate reading the wrong wire field (probability/p/value instead of the real noul), which would have abstained on every real answer — fixed. Full numbers: decision-gate-study.md.
[[gate]]
id = "verify"
decision = { provider = "typesafe", model = "jev-latest", threshold = 0.6 }
# api_key_env defaults to TYPESAFE_API_KEY; base_url defaults to https://api.typesafe.aiprior's real-data replay (OpenJev) says PROCEED, blending a learned signal into it is BLEND-NEUTRAL, hosted Jev is still unmeasured, and the decision gate measured against OpenJev is NOT-RECOMMENDED. See Related work for the full field comparison.Who it's for#
The three ideal-fit profiles
- Teams running agents at volume who pay for a frontier model on requests a cheap model could have handled — and want the savings without gambling on quality.
- Anyone who has to answer "why did the model say that?" — regulated workflows, incident reviews, or a customer dispute. The receipt is the answer.
- Builders who already have a notion of "good" — a test suite, a schema, a rubric — and want routing decided by that definition instead of a vendor's opaque classifier.
Status#
Firstpass is v0.4.0, pre-GA, and honest about the line between shipped and researched.
Shipped vs. in progress
| Shipped & verified | Next / research |
|---|---|
| Both wire dialects, structured enforce | Elastic verification (validated, phasing in) |
All six gate kinds (incl. experimental, unmeasured decision) | Per-query difficulty prediction |
| Start-rung bandit, speculation, failover | Session-level routing |
| Conformal guarantee + Learn-then-Test | Hosted plane (gated on external audit) |
| Adaptive threshold, OPE, savings/evals | crates.io publish, 30-day soak |
| Receipts + export/verify + durable mode | |
| Modes, provider-smoke CI |
escalation.elastic in your config. The conformal guarantee and the served-failure bound hold whether or not elastic mode is on.Where to go next#
- How it works — the full request flow, observe vs enforce, and the anatomy of one decision.
- Gates — the mechanism that makes routing proof-based.
- The guarantee — how the served-failure bound is computed, with the real data.
- CLI reference — every subcommand, from
onboardtoverify.