/plugin marketplace add dshakes/compass/plugin install core@compassbrew install dshakes/tap/compasscompass quickstartgemini extensions install https://github.com/dshakes/compassgit clone https://github.com/dshakes/compass ~/compass && cd ~/compass && ./quickstart.shcurl | sh Β· fully reversible Β· asks firstEveryone ships the same models. The edge is configuration you can trust β and compass hands you the number a skeptic can reproduce in 30 seconds.
Usage trackers report spend. compass enforces it β COMPASS_MAX_USD=5 hard-stops the session before the next tool call. No $40 surprise while you're away.
Catastrophic commands and secret writes are denied before they run β and the policy is eval-gated in CI against a labeled corpus, not asserted. rm -rf / β denied; rm -rf ./build β allowed.
It reviews, security-checks, tests, cross-audits with a second model, and pushes its own fixes until green. The one thing it never does: press merge. That gate is yours, permanently.
No login, no tokens. Clone the repo and run any of these β the number prints on your own machine.
Then ask the agent to rm -rf / or write a .env β denied. rm -rf ./build β allowed.
A poisoned repo or web page can't quietly turn your agent against you β and you can measure how well that holds.
Blocked, not warned β before the next tool call. A per-day cap covers unattended fleet runs.
Cheap work goes to cheap models β ~62% cheaper than all-Opus at 96.9% routing accuracy. Run compass bench to reproduce it.
The budget cap really halts before the next tool call; the loop really pushes its own fixes until tests pass. Don't take the animation's word for it β reproduce both on your machine with export COMPASS_MAX_USD=5 and /review.
Cost climbs to the cap β the next tool call is blocked, not warned. Your $5 stays $5.
It drives a change to passing tests on its own β then stops at your merge.
A one-shot agent stops at its first wrong answer. compass loops β generate β test β critique β fix β repeat against a gate β so quality comes from iteration, not one lucky prompt.
The same closed loop runs at four scales β each until a gate says done, then it stops at a human:
generate β test β critique β fix β one change driven to green.stops when tests + review pass
review β auto-fix the Blocking findings β re-review, round-capped Γ3.hands to a human if still red
the whole pipeline, scheduled across every repo you own, overnight, test-gated.a PR per repo β approve from your phone
parallel agents that fan out, fact-check each other, and converge.one synthesized answer
The whole toolkit β guardrails, the crew, the router, the CLI β is live the moment you install. The autonomous loops sit on top, opt-in.
The review β security β tests β cross-audit β auto-fix β re-review loop on every PR. Parallel reviewers, round-capped fixes, a second-model audit β then it stops at your merge.
09-sdlc βThe loop, scheduled across all your repos through a test gate β overnight. Wake up to a PR per repo you can approve from your phone.
14-fleet βA hook layer that blocks disasters, catches secrets, auto-formats, and keeps a JSONL audit log.
16-hardening βEval-gated defense vs prompt-injection, config poisoning, safety-override, MCP tool-poisoning, malware & insecure code.
17-red-team βStandalone, reusable: a deterministic keyword tier-picker (eval-gated 96.9%), with an opt-in 9-stage cost-aware engine. Measured, not vibes.
router/ βCost-tiered expert subagents, slash-commands, and dynamic workflows that fact-check each other.
roster βcompass gate β a localhost proxy that hard-caps spend for Codex, Gemini, and any SDK, not just Claude Code. The budget wall for every agent.
compass audit-plugin β vet a third-party plugin or marketplace (injection Β· tool-poisoning Β· unpinned MCP Β· fetch-and-execute hooks) before you install it.
onboard Β· impact Β· drift Β· scan Β· redteam Β· audit-plugin Β· gate Β· sandbox Β· verify Β· spend Β· dashboard
11-using βCurated, version-pinned MCP servers plus opt-in language-server intelligence.
04-mcp βClaude Code Β· Codex Β· Gemini β plus Cursor/Windsurf/Copilot via the AGENTS.md standard.
12-every-agent βEvery source outside the boundary β a pasted prompt, a fetched page, an MCP tool result, a cloned repo's CLAUDE.md β is decoded, normalized, and scored before it can steer the agent.
The cardinal rule holds underneath it all: external content is data, not instructions β and the human gate is what actually protects you.
CLAUDE.md Β· AGENTS.md Β· GEMINI.md are one file. A git pull updates them all at once.
plugin + marketplace
plugin Β· AGENTS.md
extension Β· GEMINI.md
AGENTS.md standard
All reversible, version-pinnable, no curl | sh. You need an AI assistant + git. No API keys for the guardrails, crew & CLI.
Own & edit your config β recommended. Previews every change, asks first, fully reversible.
git clone https://github.com/dshakes/compass ~/compass && cd ~/compass./quickstart.shManaged & versioned via Homebrew.
brew install dshakes/tap/compasscompass quickstartcompass quickstartNo terminal β ideal for a team. Run these inside Claude Code.
/plugin marketplace add dshakes/compass/plugin install core@compassFull control β preview, install, validate.
make dry-runmake install && make doctorAgents prepare; humans push, merge, deploy. Required checks + code-owner approval enforce it.
No curl | sh. The installer backs up what it replaces; make uninstall removes only what it added. Pin a tag, not main.
They reduce footguns. Keep least-privilege creds and review diffs. For untrusted code, compass sandbox is a real boundary.
No telemetry. Pinned MCP servers reach only the endpoints you'd expect; hooks are short, commented shell scripts you can disable.
Install once. Feel it in a minute β ask for a dangerous command (blocked), run /review on your diff, or route a task. No tokens, no signup.
git clone https://github.com/dshakes/compass ~/compass && cd ~/compass && ./quickstart.sh