Firstpass v0.3.0
DocsGuaranteeBenchmarksCLIGitHub

Docs / Configuration

Configuration

Firstpass reads a single firstpass.toml, named explicitly by FIRSTPASS_CONFIG, plus a handful of environment variables. This is the full reference — routes, ladders, gates, escalation, budgets, providers, and modes — and the multi-tenant options.

A working config is one [[route]] block. Every other section has a default and ships off — an optional feature absent from your config file is byte-invisible to callers. Add sections only when you reach for that behaviour.
Anatomy of firstpass.toml — one [[route]] block is required; every other section has a default and ships off firstpass.toml REQUIRED — A MINIMAL WORKING CONFIG [[route]] ladder ordered rungs · cheap → capable [[route]] gates non-empty/json-valid built in, no block needed OPTIONAL — EVERY SECTION BELOW HAS A DEFAULT [escalation.bandit] learned start rung · off [escalation] speculation prefetch next rung · off [budget] spend ceiling · uncapped by default [[provider]] · [[gate]] custom endpoints · custom gate kinds a running cascade ladder + gates, wired
A working config is one [[route]] block: a match, a mode, and a ladder. Every other section has a default and ships off — add it only when you reach for it.

Config shape#

One required block: [[route]], with a match, a mode, and a ladder. Everything else — custom gates, budgets, escalation tuning, providers — has a default and ships off. FIRSTPASS_CONFIG has no default path: an unset env var means routing config is inert, full stop, regardless of whether a firstpass.toml sits next to the binary.

A minimal config is one [[route]]: match the traffic, pick a mode, list the ladder. The built-in gate ids non-empty and json-valid need no extra block — only custom gate kinds (schema, subprocess, judge, consistency, decision) do. Point Firstpass at the file with FIRSTPASS_CONFIG=/path/to/firstpass.toml; there is no default path, so leaving it unset means routing runs with no config at all, not a config named firstpass.toml in the working directory.

toml
# firstpass.toml — a minimal, working config
[[route]]
match  = {}
mode   = "enforce"
ladder = ["anthropic/claude-haiku-4-5", "anthropic/claude-sonnet-5"]
gates  = ["non-empty"]
Key idea Every section not present in the file keeps its default — and the default for every optional section is off. A request through a minimal config is externally identical to a direct provider call for the traffic it does not match: same body, same model reply, same latency path. The config file is how you opt in to routing, gating, and learning; it does not change anything you do not touch.
Common mistake A hand-written firstpass.toml with a top-level [ladder] table or a [[gate]] kind = "..." field. Neither exists in the parser — the config is deny_unknown_fields, so both fail at startup with an unknown-field error rather than silently doing nothing. The ladder lives inside a [[route]] block as ladder = [...]; a gate's kind is inferred from which sub-field is set (cmd / judge / consistency / schema / decision), never a kind key.
One required block, one required env var (FIRSTPASS_CONFIG, no default path). Every optional section defaults to off and adds no behaviour until you add it.

[[route]]#

Repeatable and ordered — first match wins. Routing is a predicate over request features (calling agent, subagent, task kind, language) — not a URL path: both wire endpoints (/v1/messages, /v1/chat/completions) share the same route table. An empty match = {} matches everything, so put it last as the catch-all.

Each [[route]] binds a slice of traffic to a mode and a ladder. Routes are matched top to bottom; the first one whose match claims the request wins.

KeyTypeDefaultMeaning
matchtable{} (matches all)agent, subagent (string or list), task_kind, language (string or list) — every present field is an AND-constraint.
mode"enforce" | "observe"—, requiredGate before serving (enforce) or forward + record only (observe).
ladderlist of provider/model[]Ordered cheapest→most capable. Firstpass always opens on entry 0.
gateslist of gate ids[]Inline gates run before serving (enforce mode).
deferred_gateslist of gate ids[]Run asynchronously after serving; verdicts are a learning signal and never block the response.
routing_modemode nameNonePer-route preset; see Routing modes.
rollout[route.rollout]NonePercentage ramp; see Progressive delivery.
shadow[route.shadow]NoneOff-hot-path measurement; see Progressive delivery.
toml
[[route]]
match = { agent = "claude-code", subagent = ["test-runner", "explore"] }
mode  = "enforce"
ladder = ["anthropic/claude-haiku-4-5", "anthropic/claude-sonnet-5"]
gates  = ["non-empty", "judge"]
deferred_gates = ["consistency-check"]
routing_mode = "quality"   # optional — observe|cost|balanced|quality|latency|max

# catch-all last: an empty match claims everything above it did not
[[route]]
match = {}
mode  = "observe"
ladder = ["anthropic/claude-opus-4-8"]
Common mistake Writing match = "/v1/chat/completions" to scope a route to one wire endpoint. match is a struct of request features — it has no concept of a URL path. Both inbound endpoints share the same route table and are distinguished only by which wire dialect the caller used to talk to Firstpass, translated transparently either way. Scope routes by agent / subagent / task_kind / language instead.
Under the hood An unlisted gate id (misspelled, or never defined by a [[gate]] block) is not a parse error — it is skipped at request time with a logged warning, and the rung serves as if that gate were never named. Check the logs after any route/gate change; a silently-skipped gate looks identical to a passing one until you read them.
First matching route wins; match is feature-based, not path-based. Order specific routes before the catch-all.

Gates#

non-empty and json-valid are built in — reference them directly in gates/deferred_gates, no block needed. Every other kind (schema, subprocess, judge, consistency, decision) needs a [[gate]] block with a unique id and exactly one kind-defining field. All gates on a rung must pass. See Gates for full semantics.

A [[gate]] block has an id plus exactly one of cmd (subprocess), judge, consistency, schema, or decision — there is no kind field; the kind is inferred from which one is set, and setting zero or more than one is rejected at parse.

FieldSets kindRequired sub-fieldsMeaning
id——Unique id, referenced from a route's gates/deferred_gates.
cmdsubprocesscmd (argv list); timeout_ms optional, default 30000Any executable — candidate on stdin, JSON verdict on stdout.
judgejudge{ model, threshold?, rubric? } — threshold default 0.7A different model grades the candidate; abstains if model equals the candidate's.
consistencyconsistency{ model, k?, threshold } — k default 3, must be in [2,8]Resample k times, score by agreement. threshold required.
schemaschemaa JSON value: { type, required, properties }JSON-Schema subset: top-level type, required, per-property type.
decisiondecision{ provider = "typesafe", model?, threshold?, api_key_env?, base_url?, timeout_ms? }TypeSafe Jev cheap external verifier. See below — measured against OpenJev, not-recommended.
on_abstainall kinds"fail_open" (default) | "fail_closed"What an abstain means for serving.
toml
[[gate]]
id     = "shape"
schema = { type = "object", required = ["total", "items"] }
on_abstain = "fail_closed"

[[gate]]
id  = "unit-tests"
cmd = ["python", "run_tests.py"]
timeout_ms = 5000

[[gate]]
id    = "judge"
judge = { model = "anthropic/claude-opus-4-8", threshold = 0.7, rubric = "correct and complete" }
Under the hood Gates run after the rung responds but before that response is forwarded to the caller. A gate failure escalates to the next rung silently — the caller never sees the failing output or the gate verdict, only the final served answer. on_abstain = "fail_open" (the default) treats an inconclusive gate as a pass; "fail_closed" treats it as a failure and triggers escalation.

decision — TypeSafe Jev (experimental)

The decision gate calls TypeSafe's Jev, a cheap external decision model, and asks a single yes/no "does this response satisfy the request?" question — pass iff P(yes) >= threshold. It is orders of magnitude cheaper than a frontier LLM judge, and unlike the judge gate it is not a native LLM call through your own provider credentials.

toml
[[gate]]
id       = "decision"
decision = { provider = "typesafe", model = "jev-latest", threshold = 0.6 }
# api_key_env defaults to TYPESAFE_API_KEY; base_url defaults to https://api.typesafe.ai
NOT-RECOMMENDED (measured against OpenJev) Scored on 974 served MBPP answers (111 oracle-wrong, VRBench hidden-test oracle) with local OpenJev as the verifier: catch rate 0.2162 [0.1441, 0.2973], collateral 0.1031 [0.0834, 0.1228], AUC 0.6310 [0.5711, 0.6886] — below the pre-registered bar (catch ≥0.30 and collateral ≤0.05). Verdict: NOT-RECOMMENDED at τ=0.5; treat it as an experimental cheap pre-filter, not a drop-in replacement for judge, and do not pair it with OpenJev as a verifier. Hosted Jev as a gate remains unmeasured. Full numbers: decision-gate-study.md. Two failure paths are fail-safe by design: a timeout, transport error, non-2xx response, or an unparseable reply all resolve to Abstain, never a fabricated pass. If api_key_env names an env var that is unset at startup, the gate is skipped entirely with a logged warning — it does not block startup and does not silently pass everything; check the logs for "decision gate API key env var is unset — skipped".
Built-ins need no block; every other kind needs one field naming it, never a kind key. decision is experimental — measured NOT-RECOMMENDED against OpenJev, fail-safe to Abstain on error, skipped-with-warning if its key is missing.

[escalation]#

Escalation limits and tuning, all under one table. max_rungs_per_request (default 3) is the hard ceiling on the climb. speculation (default 0, off) fires the next rung in parallel to trade tokens for latency. Every other sub-block here is opt-in and default-off.
KeyTypeDefaultMeaning
max_rungs_per_requestint3Hard ceiling on rungs climbed within one request.
session_promotion{ after_failures, window, probe_every?, max_sessions? }NoneAfter N failures in a sliding window, start the rest of that session higher.
speculationint0 (serial)Prefetch depth: fire this many rungs ahead concurrently.
speculation_band[low, high]NoneOnly prefetch when the gate-pass estimate falls in this marginal zone.
serve_thresholdfloat [0,1]None (serve on Pass)Calibrated conformal serve threshold on aggregate gate score.
enforce_structuredbooltrueRoute tool-calling/multimodal requests through enforce.
prompt_cacheboolfalseInsert Anthropic prompt-cache breakpoints on the stable prefix.
toml
[escalation]
max_rungs_per_request = 3
session_promotion = { after_failures = 3, window = "30m" }
speculation = 0
# speculation_band = [0.3, 0.7]   # marginal zone; requires a warm bandit
# serve_threshold = 0.5           # omit to serve on a plain gate Pass

[escalation.bandit] — learned start rung

Default-off. When set, Firstpass learns which start rung tends to win for a given traffic pattern and opens there instead of always at ladder[0]. Gate evaluation and escalation are unchanged — a request that opens on rung 1 because the bandit says so still escalates to rung 2 if rung 1's gate fails (detail).

toml
[escalation.bandit]
algorithm         = "thompson"   # ucb1 (default, deterministic) | thompson (Beta, stochastic)
discount          = 0.99         # (0,1]; 1.0 = no forgetting (default)
min_observations  = 50           # cold-start guard: below this, every request starts at rung 0
exploration       = 1.0          # UCB1 constant c (Auer et al. 2002); ignored by thompson

[escalation.prior] — decision-model prior (experimental, default-off)

A pre-generation prior from TypeSafe's Jev — a per-query estimate of which ladder rung is the least capable one that fully handles the request, blended into the start-rung bandit as a Beta pseudo-observation. This is a prior, not a verifier: the gate still runs on every served attempt exactly as without it, and a confident prior can only move where the ladder starts, never what gets served (ADR 0013).

toml
[escalation.prior]
provider    = "typesafe"                # the only accepted value today
model       = "jev-latest"              # default shown
api_key_env = "TYPESAFE_API_KEY"        # default shown
base_url    = "https://api.typesafe.ai" # default shown
timeout_ms  = 150                       # default shown; a timeout fails open (no prior, no cost)
strength    = 10                        # Beta pseudo-count weight vs. real gate-verdict counts
rungs = [
  "rung 0 is the least capable tier that fully handles this request",
  "rung 1 is the least capable tier that fully handles this request",
]  # one line per ladder rung — length must match the enforce route's ladder
Simulation said STOP; a real-data replay says PROCEED The ADR 0013 pre-registered simulation gate required the prior to beat plain Firstpass on $/success at noise σ ≤ 0.2 without raising served-failure. It did not: the two arms tied at σ=0 and the prior lost beyond it — that verdict is history now. A second pre-registration replayed the same prior mechanism on 2,418 real recorded MBPP outcomes across three ladders, using OpenJev (Apache-2.0, DiffusionGemma 26B-A4B, run locally — not TypeSafe's hosted Jev) as the prior source: pooled $/success $0.01126 → $0.01075 (−4.6%), CI of the difference [-0.00074, -0.00031] excludes 0, served-failure held (0.0786 → 0.0778). Verdict: PROCEED. The win is ladder-dependent (−6.3% haiku→sonnet, −2.4% haiku→opus, 0% on the ~20x-price-ratio gpt-4.1-mini→gpt-5.5 ladder) and says nothing about hosted Jev's own accuracy. The block stays default-off pending a hosted-Jev measurement. Full numbers: openjev-prior-replay.md. If its API key env var is unset, the prior is disabled fail-open (logged warning, zero cost) rather than blocking startup — this is also true against a keyless local OpenJev, which still needs a placeholder api_key_env value (see below).
Blending a learned signal in: BLEND-NEUTRAL A third pre-registration tested whether blending traffic-learned pass rates into the prior (the proxy's own posterior-mean formula, strength pseudo-counts) beats the prior alone. Pooled (n=2418): prior+learned $0.01094 vs prior $0.01075 — paired diff +0.00019 [+0.00004, +0.00036], excludes 0, the blend is worse. Verdict: BLEND-NEUTRAL — the prior alone stays the recommendation. This run also traced an earlier "cost-aware learned-p" arm to a hindsight leak: it decided and bucketed on each task's own realized cost, which only exists after generation, so its $0.00919 pooled figure (and any earlier "~22% cheaper than first-pass" claim) is withdrawn — the honest ex-ante figure is $0.01109 pooled, only ~1.5% under first-pass. See ADR 0013. Full numbers: prior-blend-replay.md.

OpenJev (local)

Both [escalation.prior] and the decision gate speak the same POST /v1/systemone contract, so either can point at a locally-run OpenJev (Apache-2.0, DiffusionGemma 26B-A4B — razorback16/openjev) instead of TypeSafe's hosted Jev — this is what the replay above measured.

bash
# Apple Silicon, ~16 GB unified memory — binds 127.0.0.1:8080
OPENJEV_BACKEND=mlx python -m openjev
# NVIDIA, 24 GB+ VRAM:
# docker compose up
toml
[escalation.prior]
provider    = "typesafe"              # still the only accepted value — OpenJev speaks the same wire contract
base_url    = "http://127.0.0.1:8080"
api_key_env = "TYPESAFE_API_KEY"      # the client always sends a bearer token — set any placeholder, e.g. TYPESAFE_API_KEY=local

api_key_env is required even against a keyless local server — the client always attaches bearer_auth, so a missing env var disables the prior fail-open rather than sending an empty token (crates/firstpass-proxy/src/run.rs, crates/firstpass-proxy/src/gate.rs). Point the decision gate's base_url at the same server to use OpenJev as the verifier instead of the prior.

max_rungs_per_request and speculation are the two knobs most configs touch; bandit and prior are opt-in start-rung estimators layered on top — neither ever changes what a failing gate serves. prior's real-data replay (OpenJev) says PROCEED; blending a learned signal into it is BLEND-NEUTRAL; hosted Jev remains unmeasured.

[budget]#

Every cap defaults to None — uncapped — until you set it. Caps are hard ceilings enforced before any upstream call; a request that would exceed one is not forwarded.
KeyDefaultCeiling scope
per_request_usdNone (uncapped)All rungs attempted for a single request, gate cost included
per_session_usdNone (uncapped)Cumulative spend for a session
per_day_usdNone (uncapped)Daily ceiling
on_exhausted"serve_best_attempt""serve_best_attempt" (default) | "error" — what happens when a cap is hit
toml
[budget]
per_request_usd = 0.50
per_session_usd = 10.00
per_day_usd     = 250.00
on_exhausted    = "serve_best_attempt"   # serve_best_attempt (default) | error
Common mistake Speculation ([escalation] speculation) pre-pays for the next rung on every marginal call. If most calls in the speculation_band end up passing the cheap rung anyway, you are paying for two model calls on each one — the latency saving is real, but so is the token cost. Profile your gate's actual flip rate in the band before enabling.
Key idea Elastic verification (escalation.elastic) is a separate opt-in feature — config-gated and off by default. It adjusts verification intensity proportional to doubt on the serving path and requires a calibrated λ. Do not enable it without reading Elastic verification first.
Every cap is None (uncapped) until you set it — the numbers above are an example, not the shipped default. A request that would exceed any set ceiling is rejected before the upstream call, not truncated after.

[[provider]]#

Repeatable. Each block registers a model endpoint as a named provider. Reference it in a route's ladder as id/model-id. Only anthropic and openai are built in — anything else (Groq, a self-hosted vLLM/Ollama, Bedrock, Vertex) needs its own [[provider]] block.

Repeatable. Adds a provider endpoint, keyed by id (detail).

toml
[[provider]]
id          = "groq"
dialect     = "openai"             # openai | anthropic | gemini — wire format only
base_url    = "https://api.groq.com/openai"
api_key_env = "GROQ_API_KEY"       # omit entirely for a keyless local endpoint
auth        = "api_key"            # api_key (default) | aws_sigv4 | gcp_oauth
KeyDefaultMeaning
id—, requiredLadder prefix: id/model-id. Overrides a built-in of the same id.
dialect—, requiredWire format: openai | anthropic | gemini. Orthogonal to auth.
base_url""Base URL. Unused for aws_sigv4/gcp_oauth, which build the URL from region/project.
api_key_envNoneEnv var the key is read from. Omit for a keyless local endpoint (Ollama, vLLM).
auth"api_key"Credentialing scheme: api_key (header) | aws_sigv4 (Bedrock) | gcp_oauth (Vertex).
Every ladder rung needs a price A rung that names a model with no entry in the built-in price table and no matching [[price]] block is rejected at parse — there is no silent $0.00 fallback, because an unpriced rung would leave [budget] caps un-trippable and record a false cost in the audit trace. A self-hosted model you don't pay for still needs a [[price]] block declaring 0.0 explicitly.
toml
[[price]]
model           = "myllm/local-7b"
input_per_mtok  = 0.0    # USD per 1M input tokens
output_per_mtok = 0.0    # USD per 1M output tokens
One block per custom provider; reference it in a route's ladder as id/model-id. Every rung needs a price — built in or declared, even 0.0. See Providers for API key wiring.

Routing modes#

A route's routing_mode controls how aggressively Firstpass optimises within its ladder — six modes, from full passthrough (observe) to always-top (max). Set per [[route]]; see mode precedence for how a per-request override interacts with it.
toml
[[route]]
match        = { task_kind = "code_edit" }
mode         = "enforce"
ladder       = ["anthropic/claude-haiku-4-5", "anthropic/claude-sonnet-5"]
gates        = ["non-empty"]
routing_mode = "cost"        # observe|cost|balanced|quality|latency|max
ModeBehaviour
observeByte-passthrough; no routing decisions, receipts recorded only
costCheapest rung that clears the gate
balancedBalance cost against quality
qualityPrefer quality; climb earlier
latencyPrefer lowest latency
maxPin to the top rung (ladder[len-1])
routing_mode lives on the route, not a separate table; it controls how aggressively Firstpass optimises inside that route's ladder. See Routing & escalation for full semantics.

Progressive delivery#

Three default-off mechanisms that turn "flip to enforce and hope" into a path you can walk: shadow measures what Firstpass would have done, rollout ramps a slice of traffic, and the guardrail watches the served-failure bound and pulls back if it degrades. Omit them and behavior is byte-identical to before.

Enforce used to be all-or-nothing, and observe couldn't answer the question it existed for. These three compose into shadow → ramp → guard. Design rationale is in ADR 0009.

Shadow — get your own number first

On a sample of observed requests, Firstpass runs the ladder off the hot path and records what it would have served and what that would have cost. The response your caller gets is untouched — shadow work starts only after it has already gone back. It makes real model calls, which is why there is no default sample rate and the daily ceiling is required rather than optional.

toml
[route.shadow]
sample_rate     = 0.10      # fraction of observed requests scored
max_usd_per_day = 5.00      # hard ceiling; shadow stops and records that it stopped

Rollout — ramp a slice, not a guess

Enforce only part of a route's matched traffic. Bucketing is a deterministic function of a stable key, so a conversation never flips arm mid-thread. That matters for more than user experience: the published served-failure bound is computed over the population that was served, and a per-request coin flip would make that population an unstable sample of nothing. Raising percent never ejects a session that was already enforcing.

toml
[route.rollout]
percent = 5.0
key     = "session"         # session (default) | request | tenant

Guardrail — defend the number you published

Watches the same Hoeffding bound that earns the published guarantee, over a trailing window of resolved outcomes from /v1/feedback. Never gate verdicts: a guardrail scored on Firstpass's own opinion would keep enforcing exactly when that opinion had drifted.

toml
# top-level Config field — NOT nested inside [guardrail]
guardrail_cooldown_secs = 3600

[guardrail]
alpha    = 0.10             # the served-failure target being defended
delta    = 0.05             # confidence; matches the published bound
window   = 500              # trailing resolved outcomes
min_n    = 200              # never act on less evidence than this
action   = "demote"         # demote | alarm
Why demote is the default A demoted route serves exactly as it would without Firstpass. The failure mode is you stop saving money, never you keep serving worse answers. Demotion is sticky for the cooldown because re-promoting on one good window makes routing flap, and every flap is a visible change for real users. Use alarm if you want a human in the loop.
min_n is not a formality Below ln(1/delta) / (2·alpha²) the confidence slack alone exceeds alpha — so a window with zero failures would breach and every healthy route would be demoted. For alpha=0.10, delta=0.05 the floor is 150. Such a config is rejected at parse with the arithmetic shown, rather than quietly demoting everything.
Sizing min_n for your real failure rate That floor only guarantees a zero-failure window can't breach. The guardrail judges on every resolved outcome, so what governs sensitivity is the bound at min_n — where the slack is widest — not at the full window. A route genuinely running at rate r below your target stays under only once n > ln(1/delta) / (2·(alpha−r)²). At alpha=0.10, delta=0.05 a route sitting at 5% needs n > 600, well above the 150 floor — set min_n too low and a perfectly acceptable route trips the moment judging begins.
shadow → ramp → guard. Measure on your own traffic, enforce a slice, and let the bound you published defend itself. All three off by default.

Environment variables#

Seven variables control listen address, config file, and operating mode. Start with FIRSTPASS_MODE=observe — it records receipts and passes every request through unchanged. FIRSTPASS_CONFIG has no default path — leave it unset and Firstpass runs with no routing config at all, even if a firstpass.toml sits right next to the binary.
VariableValuesMeaning
FIRSTPASS_MODEobserve | enforceGlobal default mode (default observe). A route's own mode still governs whether it gates.
FIRSTPASS_CONFIGpathNo default. Unset ⇒ no firstpass.toml is read, regardless of what's in the working directory.
FIRSTPASS_DBpathReceipt / state store location (default firstpass.db).
FIRSTPASS_BINDhost:portListen address (default 127.0.0.1:8080).
FIRSTPASS_RECEIPTSbest_effort | durableNever-drop receipts (default best_effort; detail).
FIRSTPASS_MODE_PROFILEmode nameEnv-level default mode profile (default balanced).
FIRSTPASS_MAX_CONCURRENCYintMax in-flight upstream requests (default 512).
Common mistake Writing a valid firstpass.toml in the project root and assuming Firstpass picks it up. It doesn't — FIRSTPASS_CONFIG is the only thing that makes a config file live; there is no implicit ./firstpass.toml lookup. A written-but-unreferenced file is completely inert.
Under the hood FIRSTPASS_MODE=observe is a byte-passthrough: every request is forwarded verbatim to the upstream, receipts are written, but no routing or gate decisions are made. The model and the caller see no difference from a direct connection — it is the zero-risk way to instrument an existing agent. Switch to enforce (globally via the env var, or per-route via mode = "enforce") once you have validated your ladder and gate config against observe-mode receipts.
FIRSTPASS_MODE=observe is shadow mode — zero behaviour change, full receipts. FIRSTPASS_CONFIG must be set explicitly; there is no default path.

Multi-tenant#

Opt-in and default-off. When enabled: Argon2id keyed auth, AES-256-GCM key custody, per-tenant rate limits. Not GA — pending external security review. Do not use for tenant isolation in production yet.

Multi-tenancy is opt-in and default-off. When enabled it provides Argon2id keyed auth, AES-256-GCM key custody, and per-tenant rate limits.

Pre-external-review The multi-tenant path is implemented but not GA — it is pending external security review. Do not rely on it for tenant isolation in production until that review lands. See ADR 0004.
Multi-tenancy is off until you configure it and not GA until the external security review lands. Do not rely on it for isolation in production today.