Docs / Configuration
Configuration
Firstpass reads a single firstpass.toml, named explicitly by FIRSTPASS_CONFIG, plus a handful of environment variables. This is the full reference — routes, ladders, gates, escalation, budgets, providers, and modes — and the multi-tenant options.
[[route]] block. Every other section has a default and ships off — an optional feature absent from your config file is byte-invisible to callers. Add sections only when you reach for that behaviour.[[route]] block: a match, a mode, and a ladder. Every other section has a default and ships off — add it only when you reach for it.Config shape#
[[route]], with a match, a mode, and a ladder. Everything else — custom gates, budgets, escalation tuning, providers — has a default and ships off. FIRSTPASS_CONFIG has no default path: an unset env var means routing config is inert, full stop, regardless of whether a firstpass.toml sits next to the binary.A minimal config is one [[route]]: match the traffic, pick a mode, list the ladder. The built-in gate ids non-empty and json-valid need no extra block — only custom gate kinds (schema, subprocess, judge, consistency, decision) do. Point Firstpass at the file with FIRSTPASS_CONFIG=/path/to/firstpass.toml; there is no default path, so leaving it unset means routing runs with no config at all, not a config named firstpass.toml in the working directory.
# firstpass.toml — a minimal, working config
[[route]]
match = {}
mode = "enforce"
ladder = ["anthropic/claude-haiku-4-5", "anthropic/claude-sonnet-5"]
gates = ["non-empty"]firstpass.toml with a top-level [ladder] table or a [[gate]] kind = "..." field. Neither exists in the parser — the config is deny_unknown_fields, so both fail at startup with an unknown-field error rather than silently doing nothing. The ladder lives inside a [[route]] block as ladder = [...]; a gate's kind is inferred from which sub-field is set (cmd / judge / consistency / schema / decision), never a kind key.FIRSTPASS_CONFIG, no default path). Every optional section defaults to off and adds no behaviour until you add it.[[route]]#
/v1/messages, /v1/chat/completions) share the same route table. An empty match = {} matches everything, so put it last as the catch-all.Each [[route]] binds a slice of traffic to a mode and a ladder. Routes are matched top to bottom; the first one whose match claims the request wins.
| Key | Type | Default | Meaning |
|---|---|---|---|
match | table | {} (matches all) | agent, subagent (string or list), task_kind, language (string or list) — every present field is an AND-constraint. |
mode | "enforce" | "observe" | —, required | Gate before serving (enforce) or forward + record only (observe). |
ladder | list of provider/model | [] | Ordered cheapest→most capable. Firstpass always opens on entry 0. |
gates | list of gate ids | [] | Inline gates run before serving (enforce mode). |
deferred_gates | list of gate ids | [] | Run asynchronously after serving; verdicts are a learning signal and never block the response. |
routing_mode | mode name | None | Per-route preset; see Routing modes. |
rollout | [route.rollout] | None | Percentage ramp; see Progressive delivery. |
shadow | [route.shadow] | None | Off-hot-path measurement; see Progressive delivery. |
[[route]]
match = { agent = "claude-code", subagent = ["test-runner", "explore"] }
mode = "enforce"
ladder = ["anthropic/claude-haiku-4-5", "anthropic/claude-sonnet-5"]
gates = ["non-empty", "judge"]
deferred_gates = ["consistency-check"]
routing_mode = "quality" # optional — observe|cost|balanced|quality|latency|max
# catch-all last: an empty match claims everything above it did not
[[route]]
match = {}
mode = "observe"
ladder = ["anthropic/claude-opus-4-8"]match = "/v1/chat/completions" to scope a route to one wire endpoint. match is a struct of request features — it has no concept of a URL path. Both inbound endpoints share the same route table and are distinguished only by which wire dialect the caller used to talk to Firstpass, translated transparently either way. Scope routes by agent / subagent / task_kind / language instead.[[gate]] block) is not a parse error — it is skipped at request time with a logged warning, and the rung serves as if that gate were never named. Check the logs after any route/gate change; a silently-skipped gate looks identical to a passing one until you read them.match is feature-based, not path-based. Order specific routes before the catch-all.Gates#
non-empty and json-valid are built in — reference them directly in gates/deferred_gates, no block needed. Every other kind (schema, subprocess, judge, consistency, decision) needs a [[gate]] block with a unique id and exactly one kind-defining field. All gates on a rung must pass. See Gates for full semantics.A [[gate]] block has an id plus exactly one of cmd (subprocess), judge, consistency, schema, or decision — there is no kind field; the kind is inferred from which one is set, and setting zero or more than one is rejected at parse.
| Field | Sets kind | Required sub-fields | Meaning |
|---|---|---|---|
id | — | — | Unique id, referenced from a route's gates/deferred_gates. |
cmd | subprocess | cmd (argv list); timeout_ms optional, default 30000 | Any executable — candidate on stdin, JSON verdict on stdout. |
judge | judge | { model, threshold?, rubric? } — threshold default 0.7 | A different model grades the candidate; abstains if model equals the candidate's. |
consistency | consistency | { model, k?, threshold } — k default 3, must be in [2,8] | Resample k times, score by agreement. threshold required. |
schema | schema | a JSON value: { type, required, properties } | JSON-Schema subset: top-level type, required, per-property type. |
decision | decision | { provider = "typesafe", model?, threshold?, api_key_env?, base_url?, timeout_ms? } | TypeSafe Jev cheap external verifier. See below — measured against OpenJev, not-recommended. |
on_abstain | all kinds | "fail_open" (default) | "fail_closed" | What an abstain means for serving. |
[[gate]]
id = "shape"
schema = { type = "object", required = ["total", "items"] }
on_abstain = "fail_closed"
[[gate]]
id = "unit-tests"
cmd = ["python", "run_tests.py"]
timeout_ms = 5000
[[gate]]
id = "judge"
judge = { model = "anthropic/claude-opus-4-8", threshold = 0.7, rubric = "correct and complete" }on_abstain = "fail_open" (the default) treats an inconclusive gate as a pass; "fail_closed" treats it as a failure and triggers escalation.decision — TypeSafe Jev (experimental)
The decision gate calls TypeSafe's Jev, a cheap external decision model, and asks a single yes/no "does this response satisfy the request?" question — pass iff P(yes) >= threshold. It is orders of magnitude cheaper than a frontier LLM judge, and unlike the judge gate it is not a native LLM call through your own provider credentials.
[[gate]]
id = "decision"
decision = { provider = "typesafe", model = "jev-latest", threshold = 0.6 }
# api_key_env defaults to TYPESAFE_API_KEY; base_url defaults to https://api.typesafe.aijudge, and do not pair it with OpenJev as a verifier. Hosted Jev as a gate remains unmeasured. Full numbers: decision-gate-study.md. Two failure paths are fail-safe by design: a timeout, transport error, non-2xx response, or an unparseable reply all resolve to Abstain, never a fabricated pass. If api_key_env names an env var that is unset at startup, the gate is skipped entirely with a logged warning — it does not block startup and does not silently pass everything; check the logs for "decision gate API key env var is unset — skipped".kind key. decision is experimental — measured NOT-RECOMMENDED against OpenJev, fail-safe to Abstain on error, skipped-with-warning if its key is missing.[escalation]#
max_rungs_per_request (default 3) is the hard ceiling on the climb. speculation (default 0, off) fires the next rung in parallel to trade tokens for latency. Every other sub-block here is opt-in and default-off.| Key | Type | Default | Meaning |
|---|---|---|---|
max_rungs_per_request | int | 3 | Hard ceiling on rungs climbed within one request. |
session_promotion | { after_failures, window, probe_every?, max_sessions? } | None | After N failures in a sliding window, start the rest of that session higher. |
speculation | int | 0 (serial) | Prefetch depth: fire this many rungs ahead concurrently. |
speculation_band | [low, high] | None | Only prefetch when the gate-pass estimate falls in this marginal zone. |
serve_threshold | float [0,1] | None (serve on Pass) | Calibrated conformal serve threshold on aggregate gate score. |
enforce_structured | bool | true | Route tool-calling/multimodal requests through enforce. |
prompt_cache | bool | false | Insert Anthropic prompt-cache breakpoints on the stable prefix. |
[escalation]
max_rungs_per_request = 3
session_promotion = { after_failures = 3, window = "30m" }
speculation = 0
# speculation_band = [0.3, 0.7] # marginal zone; requires a warm bandit
# serve_threshold = 0.5 # omit to serve on a plain gate Pass[escalation.bandit] — learned start rung
Default-off. When set, Firstpass learns which start rung tends to win for a given traffic pattern and opens there instead of always at ladder[0]. Gate evaluation and escalation are unchanged — a request that opens on rung 1 because the bandit says so still escalates to rung 2 if rung 1's gate fails (detail).
[escalation.bandit]
algorithm = "thompson" # ucb1 (default, deterministic) | thompson (Beta, stochastic)
discount = 0.99 # (0,1]; 1.0 = no forgetting (default)
min_observations = 50 # cold-start guard: below this, every request starts at rung 0
exploration = 1.0 # UCB1 constant c (Auer et al. 2002); ignored by thompson[escalation.prior] — decision-model prior (experimental, default-off)
A pre-generation prior from TypeSafe's Jev — a per-query estimate of which ladder rung is the least capable one that fully handles the request, blended into the start-rung bandit as a Beta pseudo-observation. This is a prior, not a verifier: the gate still runs on every served attempt exactly as without it, and a confident prior can only move where the ladder starts, never what gets served (ADR 0013).
[escalation.prior]
provider = "typesafe" # the only accepted value today
model = "jev-latest" # default shown
api_key_env = "TYPESAFE_API_KEY" # default shown
base_url = "https://api.typesafe.ai" # default shown
timeout_ms = 150 # default shown; a timeout fails open (no prior, no cost)
strength = 10 # Beta pseudo-count weight vs. real gate-verdict counts
rungs = [
"rung 0 is the least capable tier that fully handles this request",
"rung 1 is the least capable tier that fully handles this request",
] # one line per ladder rung — length must match the enforce route's ladder[-0.00074, -0.00031] excludes 0, served-failure held (0.0786 → 0.0778). Verdict: PROCEED. The win is ladder-dependent (−6.3% haiku→sonnet, −2.4% haiku→opus, 0% on the ~20x-price-ratio gpt-4.1-mini→gpt-5.5 ladder) and says nothing about hosted Jev's own accuracy. The block stays default-off pending a hosted-Jev measurement. Full numbers: openjev-prior-replay.md. If its API key env var is unset, the prior is disabled fail-open (logged warning, zero cost) rather than blocking startup — this is also true against a keyless local OpenJev, which still needs a placeholder api_key_env value (see below).strength pseudo-counts) beats the prior alone. Pooled (n=2418): prior+learned $0.01094 vs prior $0.01075 — paired diff +0.00019 [+0.00004, +0.00036], excludes 0, the blend is worse. Verdict: BLEND-NEUTRAL — the prior alone stays the recommendation. This run also traced an earlier "cost-aware learned-p" arm to a hindsight leak: it decided and bucketed on each task's own realized cost, which only exists after generation, so its $0.00919 pooled figure (and any earlier "~22% cheaper than first-pass" claim) is withdrawn — the honest ex-ante figure is $0.01109 pooled, only ~1.5% under first-pass. See ADR 0013. Full numbers: prior-blend-replay.md.OpenJev (local)
Both [escalation.prior] and the decision gate speak the same POST /v1/systemone contract, so either can point at a locally-run OpenJev (Apache-2.0, DiffusionGemma 26B-A4B — razorback16/openjev) instead of TypeSafe's hosted Jev — this is what the replay above measured.
# Apple Silicon, ~16 GB unified memory — binds 127.0.0.1:8080
OPENJEV_BACKEND=mlx python -m openjev
# NVIDIA, 24 GB+ VRAM:
# docker compose up[escalation.prior]
provider = "typesafe" # still the only accepted value — OpenJev speaks the same wire contract
base_url = "http://127.0.0.1:8080"
api_key_env = "TYPESAFE_API_KEY" # the client always sends a bearer token — set any placeholder, e.g. TYPESAFE_API_KEY=localapi_key_env is required even against a keyless local server — the client always attaches bearer_auth, so a missing env var disables the prior fail-open rather than sending an empty token (crates/firstpass-proxy/src/run.rs, crates/firstpass-proxy/src/gate.rs). Point the decision gate's base_url at the same server to use OpenJev as the verifier instead of the prior.
max_rungs_per_request and speculation are the two knobs most configs touch; bandit and prior are opt-in start-rung estimators layered on top — neither ever changes what a failing gate serves. prior's real-data replay (OpenJev) says PROCEED; blending a learned signal into it is BLEND-NEUTRAL; hosted Jev remains unmeasured.[budget]#
None — uncapped — until you set it. Caps are hard ceilings enforced before any upstream call; a request that would exceed one is not forwarded.| Key | Default | Ceiling scope |
|---|---|---|
per_request_usd | None (uncapped) | All rungs attempted for a single request, gate cost included |
per_session_usd | None (uncapped) | Cumulative spend for a session |
per_day_usd | None (uncapped) | Daily ceiling |
on_exhausted | "serve_best_attempt" | "serve_best_attempt" (default) | "error" — what happens when a cap is hit |
[budget]
per_request_usd = 0.50
per_session_usd = 10.00
per_day_usd = 250.00
on_exhausted = "serve_best_attempt" # serve_best_attempt (default) | error[escalation] speculation) pre-pays for the next rung on every marginal call. If most calls in the speculation_band end up passing the cheap rung anyway, you are paying for two model calls on each one — the latency saving is real, but so is the token cost. Profile your gate's actual flip rate in the band before enabling.escalation.elastic) is a separate opt-in feature — config-gated and off by default. It adjusts verification intensity proportional to doubt on the serving path and requires a calibrated λ. Do not enable it without reading Elastic verification first.None (uncapped) until you set it — the numbers above are an example, not the shipped default. A request that would exceed any set ceiling is rejected before the upstream call, not truncated after.[[provider]]#
id/model-id. Only anthropic and openai are built in — anything else (Groq, a self-hosted vLLM/Ollama, Bedrock, Vertex) needs its own [[provider]] block.Repeatable. Adds a provider endpoint, keyed by id (detail).
[[provider]]
id = "groq"
dialect = "openai" # openai | anthropic | gemini — wire format only
base_url = "https://api.groq.com/openai"
api_key_env = "GROQ_API_KEY" # omit entirely for a keyless local endpoint
auth = "api_key" # api_key (default) | aws_sigv4 | gcp_oauth| Key | Default | Meaning |
|---|---|---|
id | —, required | Ladder prefix: id/model-id. Overrides a built-in of the same id. |
dialect | —, required | Wire format: openai | anthropic | gemini. Orthogonal to auth. |
base_url | "" | Base URL. Unused for aws_sigv4/gcp_oauth, which build the URL from region/project. |
api_key_env | None | Env var the key is read from. Omit for a keyless local endpoint (Ollama, vLLM). |
auth | "api_key" | Credentialing scheme: api_key (header) | aws_sigv4 (Bedrock) | gcp_oauth (Vertex). |
[[price]] block is rejected at parse — there is no silent $0.00 fallback, because an unpriced rung would leave [budget] caps un-trippable and record a false cost in the audit trace. A self-hosted model you don't pay for still needs a [[price]] block declaring 0.0 explicitly.[[price]]
model = "myllm/local-7b"
input_per_mtok = 0.0 # USD per 1M input tokens
output_per_mtok = 0.0 # USD per 1M output tokensid/model-id. Every rung needs a price — built in or declared, even 0.0. See Providers for API key wiring.Routing modes#
routing_mode controls how aggressively Firstpass optimises within its ladder — six modes, from full passthrough (observe) to always-top (max). Set per [[route]]; see mode precedence for how a per-request override interacts with it.[[route]]
match = { task_kind = "code_edit" }
mode = "enforce"
ladder = ["anthropic/claude-haiku-4-5", "anthropic/claude-sonnet-5"]
gates = ["non-empty"]
routing_mode = "cost" # observe|cost|balanced|quality|latency|max| Mode | Behaviour |
|---|---|
observe | Byte-passthrough; no routing decisions, receipts recorded only |
cost | Cheapest rung that clears the gate |
balanced | Balance cost against quality |
quality | Prefer quality; climb earlier |
latency | Prefer lowest latency |
max | Pin to the top rung (ladder[len-1]) |
routing_mode lives on the route, not a separate table; it controls how aggressively Firstpass optimises inside that route's ladder. See Routing & escalation for full semantics.Progressive delivery#
Enforce used to be all-or-nothing, and observe couldn't answer the question it existed for. These three compose into shadow → ramp → guard. Design rationale is in ADR 0009.
Shadow — get your own number first
On a sample of observed requests, Firstpass runs the ladder off the hot path and records what it would have served and what that would have cost. The response your caller gets is untouched — shadow work starts only after it has already gone back. It makes real model calls, which is why there is no default sample rate and the daily ceiling is required rather than optional.
[route.shadow]
sample_rate = 0.10 # fraction of observed requests scored
max_usd_per_day = 5.00 # hard ceiling; shadow stops and records that it stoppedRollout — ramp a slice, not a guess
Enforce only part of a route's matched traffic. Bucketing is a deterministic function of a stable key, so a conversation never flips arm mid-thread. That matters for more than user experience: the published served-failure bound is computed over the population that was served, and a per-request coin flip would make that population an unstable sample of nothing. Raising percent never ejects a session that was already enforcing.
[route.rollout]
percent = 5.0
key = "session" # session (default) | request | tenantGuardrail — defend the number you published
Watches the same Hoeffding bound that earns the published guarantee, over a trailing window of resolved outcomes from /v1/feedback. Never gate verdicts: a guardrail scored on Firstpass's own opinion would keep enforcing exactly when that opinion had drifted.
# top-level Config field — NOT nested inside [guardrail]
guardrail_cooldown_secs = 3600
[guardrail]
alpha = 0.10 # the served-failure target being defended
delta = 0.05 # confidence; matches the published bound
window = 500 # trailing resolved outcomes
min_n = 200 # never act on less evidence than this
action = "demote" # demote | alarmalarm if you want a human in the loop.ln(1/delta) / (2·alpha²) the confidence slack alone exceeds alpha — so a window with zero failures would breach and every healthy route would be demoted. For alpha=0.10, delta=0.05 the floor is 150. Such a config is rejected at parse with the arithmetic shown, rather than quietly demoting everything.min_n — where the slack is widest — not at the full window. A route genuinely running at rate r below your target stays under only once n > ln(1/delta) / (2·(alpha−r)²). At alpha=0.10, delta=0.05 a route sitting at 5% needs n > 600, well above the 150 floor — set min_n too low and a perfectly acceptable route trips the moment judging begins.Environment variables#
FIRSTPASS_MODE=observe — it records receipts and passes every request through unchanged. FIRSTPASS_CONFIG has no default path — leave it unset and Firstpass runs with no routing config at all, even if a firstpass.toml sits right next to the binary.| Variable | Values | Meaning |
|---|---|---|
FIRSTPASS_MODE | observe | enforce | Global default mode (default observe). A route's own mode still governs whether it gates. |
FIRSTPASS_CONFIG | path | No default. Unset ⇒ no firstpass.toml is read, regardless of what's in the working directory. |
FIRSTPASS_DB | path | Receipt / state store location (default firstpass.db). |
FIRSTPASS_BIND | host:port | Listen address (default 127.0.0.1:8080). |
FIRSTPASS_RECEIPTS | best_effort | durable | Never-drop receipts (default best_effort; detail). |
FIRSTPASS_MODE_PROFILE | mode name | Env-level default mode profile (default balanced). |
FIRSTPASS_MAX_CONCURRENCY | int | Max in-flight upstream requests (default 512). |
firstpass.toml in the project root and assuming Firstpass picks it up. It doesn't — FIRSTPASS_CONFIG is the only thing that makes a config file live; there is no implicit ./firstpass.toml lookup. A written-but-unreferenced file is completely inert.FIRSTPASS_MODE=observe is a byte-passthrough: every request is forwarded verbatim to the upstream, receipts are written, but no routing or gate decisions are made. The model and the caller see no difference from a direct connection — it is the zero-risk way to instrument an existing agent. Switch to enforce (globally via the env var, or per-route via mode = "enforce") once you have validated your ladder and gate config against observe-mode receipts.FIRSTPASS_MODE=observe is shadow mode — zero behaviour change, full receipts. FIRSTPASS_CONFIG must be set explicitly; there is no default path.Multi-tenant#
Multi-tenancy is opt-in and default-off. When enabled it provides Argon2id keyed auth, AES-256-GCM key custody, and per-tenant rate limits.