Firstpass v0.4.0
DocsGuaranteeBenchmarksCLIGitHub

Docs / Overview

Overview

Firstpass is an open-source adaptive LLM router. Every request opens on the cheapest model in a ladder; a gate checks the actual output; it serves the first output that passes and escalates one rung only on gate failure. Every decision gets a tamper-evident, hash-chained receipt.

Firstpass is an open-source verification layer for LLM serving, v0.4.0, pre-GA. Your own check runs on the real output of every request and only what passes is served; routing is the consequence, not the mechanism. Every request opens on the cheapest model; a gate proves the real output before serving it; on gate failure, it escalates exactly one rung. On 974 real coding tasks that recovered +15.2 points of quality over the cheap model alone and made zero of those tasks worse (artifacts). Four moves — Route → Prove → Escalate → Learn — on every request. It sits as an HTTP reverse proxy between your agent and any provider; your agent code doesn't change.
≤10%
served-failure, 95% conf.
82%
served from cheap tier
974
real tasks in the live bound
v0.4.0
pre-GA, open source

What Firstpass is#

Most routers predict quality before any output exists — and when the prediction is wrong, you find out in production with no artifact explaining why. Firstpass proves quality after: open the cheapest model, gate the actual output, serve on pass, escalate one rung on fail. The positioning in one line: proof over prediction.

Prediction-based routing and its limits

Most model gateways route by prediction: a classifier reads the prompt and guesses which model can handle it, before a single token of output exists. If the guess is wrong, you find out in production — and there's no artifact explaining why the request went where it did.

How proof routing works

Firstpass routes by proof. It opens on the cheapest model, then gates the real output — runs your tests, checks a schema, asks a judge, or measures self-consistency. If the output passes, it serves. If it fails, it escalates exactly one rung and gates again. The cheap model handles most traffic; the expensive model is spent only when the cheap one is provably not enough.

The positioning is one line: proof over prediction — verified cascade routing.

Key idea The "proof" is your definition of good — a test suite, a JSON schema, a pass/fail judge, or self-consistency across samples. Firstpass supplies the escalation machinery and the audit; you supply the criterion. That means routing is decided by the thing you already trust, not an opaque classifier — and every decision is explained by a gate result on the actual output, captured in a hash-chained receipt.
Honest novelty The cascade mechanism itself is prior art. Firstpass is not claiming to have invented cheap-model-first escalation. What is new is the combination in the section below — the audit, the calibration, and the guarantee wrapped around the cascade.
Firstpass routes by proof — the cheap model attempts first, the gate decides from the real output, and only a provably failed gate triggers escalation.

The four moves#

Route opens the cheapest rung — no per-prompt prediction. Prove runs your gate on the real output. Escalate goes one rung on gate failure, budget-capped. Learn feeds outcomes back to self-tune the serve threshold. All four run on every request, in order, and your agent sees a single round-trip.

Everything the router does reduces to four moves, in order, on every request.

The four moves: Route opens on the cheapest model, Prove gates the real output, Escalate goes one rung up only on failure, and Learn feeds outcomes back so the serve threshold self-tunes. 1 Route open the cheapest rung in the ladder 2 Prove gate the real output, not the prompt 3 Escalate one rung on fail, budget-capped 4 Learn outcomes tune the serve threshold

What each move does

Route
Open the cheapest rung in the ladder. No per-prompt prediction — the cheap model takes the first pass.
Prove
Gate the real output, not the prompt. Built-in checks, your test suite, an LLM judge, or self-consistency.
Escalate
One rung up on gate failure — budget-capped, with cross-provider failover on transport or 5xx errors.
Learn
Deferred outcomes feed back through /v1/feedback; adaptive conformal moves the serve threshold under drift.

How it works follows a single request through all four in detail.

Under the hood Firstpass is an HTTP reverse proxy sitting below your agent. It accepts the same POST /v1/messages (Anthropic), POST /v1/chat/completions (OpenAI), or POST /v1/responses (OpenAI Responses — including tool calls, translated both ways) your agent would send to a provider, swaps only the model field when routing, and forwards everything else verbatim. SSE streaming is buffered so the gate can run, then re-emitted as SSE — your agent sees a normal response. The four moves all happen in that buffer window, transparently.
Every request follows Route → Prove → Escalate → Learn in a single round-trip; your agent points at Firstpass the same way it points at any provider.

What is actually new#

The cascade idea — cheap model first, escalate on failure — is prior art. What Firstpass adds: tamper-evident hash-chained audit, outcome-calibrated serve threshold, zero-retrain model onboarding, predict-to-start + verify-to-serve, and a distribution-free served-failure guarantee earned from a real run.

Five things, and Firstpass is precise about which are novel versus assembled from known parts:

The five novelties, precisely stated

  1. Tamper-evident audit + independent verification. Every decision is a hash-chained receipt; an external auditor re-derives the whole chain from genesis with no access to the proxy or its database. See Receipts & audit.
  2. Outcome-feedback calibration of the serve threshold. Real outcomes — not prompt features — move the threshold that decides serve-versus-escalate. See Adaptive conformal.
  3. Zero-retrain model onboarding. Add a rung as provider/model; there is no classifier to retrain when models change. See Providers.
  4. Predict-to-start, verify-to-serve. A learned start-rung bandit predicts where to begin; the gate still proves what to serve. Prediction sets the entry point; proof sets the exit. See Learned start-rung.
  5. A served-failure guarantee. A distribution-free bound on how often a wrong answer is served, earned from a real run. See The guarantee.
Key idea "Prediction sets the entry point; proof sets the exit." A learned bandit can pick a smarter starting rung for efficiency — but the gate is still what decides whether that output is served. Speed and correctness are separate problems, solved by separate mechanisms, and neither undermines the other.
The five novelties are separable and stackable — audit, calibration, and the guarantee each add value independently and compound when combined.

Decision-model routers (Jev) vs Firstpass#

TypeSafe's Jev answers one cheap closed-form question per query and routes by picking a model tier directly, then serves that tier's output unverified. Firstpass's thesis is narrower: predict the start, verify what's served — the same cheap signal can pick where the ladder opens, but the gate still runs on every attempt. [escalation.prior] is experimental, default-off: the original sim verdict was PRIOR=STOP, a second pre-registration replayed on real MBPP data with OpenJev now reads PROCEED (see below); a third found blending a learned signal into it is BLEND-NEUTRAL; the decision gate measured against OpenJev is NOT-RECOMMENDED (hosted Jev still unmeasured).

A new class of "System One" decision models — TypeSafe's Jev — answers one cheap closed-form question per query ($0.042/M input tokens, no text generation) and routes a query to a model tier by asking it to pick the tier directly. Products built on it (jev-router, prismhq/jev-router) then serve that tier's output unverified.

Firstpass takes the same cheap signal and puts it somewhere narrower: the existing start-rung bandit already had a slot for a per-query prior on where the ladder opens. The gate still runs on every attempt — a wrong prediction costs money or latency, never a wrong answer shipped.

The measured gap (simulation, not live)

Firstpass's own σ-sweep (cargo run -p firstpass-bench, n=500) puts a Jev-style unverified router's served-failure at 50.6–56.2% across the noise levels tested, against Firstpass's own gated served-failure of 15.8% on the same suite. Full numbers and both addenda: ADR 0013.

Status: simulation said STOP, real-data replay says PROCEED The original pre-registered simulation (synthetic σ-sweep) found the fused arm tying plain Firstpass at σ=0 and losing at higher noise — that's history now. A second pre-registration replayed the same prior mechanism on 2,418 real recorded MBPP outcomes across three ladders, using OpenJev (Apache-2.0, run locally — not TypeSafe's hosted Jev) as the prior source: pooled $/success $0.01126 → $0.01075 (−4.6%), CI of the difference [-0.00074, -0.00031] excludes 0, served-failure held (0.0786 → 0.0778). Verdict: PROCEED. The win is ladder-dependent (−6.3% haiku→sonnet, −2.4% haiku→opus, 0% on the ~20x-price-ratio gpt-4.1-mini→gpt-5.5 ladder) and OpenJev says nothing about hosted Jev's own accuracy. [escalation.prior] stays default-off — the replay is evidence about the prior mechanism, not about the hosted vendor. Full numbers: openjev-prior-replay.md.
Blending a learned signal in: BLEND-NEUTRAL A third pre-registration tested whether blending traffic-learned pass rates into the prior beats the prior alone. Pooled (n=2418): prior+learned $0.01094 vs prior $0.01075 — paired diff +0.00019 [+0.00004, +0.00036], excludes 0, the blend is worse. Verdict: BLEND-NEUTRAL — the prior alone stays the recommendation. This run also traced an earlier "cost-aware learned-p" arm to a hindsight leak: it decided on each task's own realized cost, which only exists after generation, so its savings claim (including any earlier "~22% cheaper than first-pass" figure) is withdrawn — see ADR 0013. Full numbers: prior-blend-replay.md.

[escalation.prior] — enabling it anyway

toml
# experimental, default-off — sim was PRIOR=STOP, real-data replay (OpenJev) says PROCEED
[escalation.prior]
provider    = "typesafe"                # the only accepted value today
base_url    = "https://api.typesafe.ai" # default; self-hosted/mock Jev overrides this
model       = "jev-latest"
api_key_env = "TYPESAFE_API_KEY"        # env var name only — the key itself is never logged
timeout_ms  = 150                       # any error fails open — no prior, no decision_prior field
strength    = 10                        # Beta pseudo-count weight of the prior
rungs = [
  "rung 0 (claude-haiku) is the least capable tier that fully handles this request",
  "rung 1 (claude-sonnet) is the least capable tier that fully handles this request",
]
Key idea rungs must have exactly one entry per rung of every enforce-mode route's ladder — a length mismatch is rejected at config parse, never silently truncated. The gate still verifies every served output; the prior only ever moves where the ladder starts.

decision gate — NOT-RECOMMENDED (OpenJev)

The same Jev model can also sit behind a [[gate]] block as a cheap external verifier (~$0.042/M input tokens) instead of a frontier LLM judge. A transport error, timeout, or unparseable reply ABSTAINs — never a fabricated pass. Measured on 974 served MBPP answers (111 oracle-wrong) with local OpenJev as the verifier: catch rate 0.2162 [0.1441, 0.2973], collateral 0.1031 [0.0834, 0.1228], AUC 0.6310 [0.5711, 0.6886] — below the pre-registered bar (catch ≥0.30 and collateral ≤0.05). Verdict: NOT-RECOMMENDED at τ=0.5. This measures OpenJev, not hosted Jev, which remains unmeasured. It also caught the gate reading the wrong wire field (probability/p/value instead of the real noul), which would have abstained on every real answer — fixed. Full numbers: decision-gate-study.md.

toml
[[gate]]
id       = "verify"
decision = { provider = "typesafe", model = "jev-latest", threshold = 0.6 }
# api_key_env defaults to TYPESAFE_API_KEY; base_url defaults to https://api.typesafe.ai
Jev-style routers guess and serve unverified; Firstpass's prior only ever moves the start rung, and the gate still decides what ships. Both new knobs are opt-in, off by default, and honestly labeled: prior's real-data replay (OpenJev) says PROCEED, blending a learned signal into it is BLEND-NEUTRAL, hosted Jev is still unmeasured, and the decision gate measured against OpenJev is NOT-RECOMMENDED. See Related work for the full field comparison.

Who it's for#

Three ideal-fit profiles: teams paying frontier prices on tasks a cheap model could handle; anyone who needs a tamper-evident "why" for every routing decision; builders with an existing definition of "good" — tests, schemas, rubrics — who want routing by their criterion instead of a vendor classifier.

The three ideal-fit profiles

Key idea If you already have a test suite that runs in seconds, you already have a gate. Plug it in and every routing decision becomes: "did the cheap model pass your tests?" The receipt that explains the routing decision comes for free — the same artifact that routes also audits.
Firstpass fits best when you have a definition of "good" you can run, and a cost problem you can see in the bill.

Status#

v0.4.0 is shipped and running. The "Shipped & verified" column is available and on by default. The "Next / research" column is gated on further validation — nothing there ships as a default, and every feature is opt-in only.

Firstpass is v0.4.0, pre-GA, and honest about the line between shipped and researched.

Shipped vs. in progress

Shipped & verifiedNext / research
Both wire dialects, structured enforceElastic verification (validated, phasing in)
All six gate kinds (incl. experimental, unmeasured decision)Per-query difficulty prediction
Start-rung bandit, speculation, failoverSession-level routing
Conformal guarantee + Learn-then-TestHosted plane (gated on external audit)
Adaptive threshold, OPE, savings/evalscrates.io publish, 30-day soak
Receipts + export/verify + durable mode
Modes, provider-smoke CI
Under the hood Elastic verification (ADR 0008) is implemented and config-gated but off by default. Enable it by supplying a calibrated λ under escalation.elastic in your config. The conformal guarantee and the served-failure bound hold whether or not elastic mode is on.
The shipped column is available today; the next column is gated on further validation — see Benchmarks for the live data behind the guarantee.

Where to go next#

New here: start with How it works for the full request lifecycle. Evaluating routing quality: go to The guarantee for the live data. Ready to build: the Quickstart gets you routed in minutes.