# RealityRouter > One router. Every model. Always the cheapest that works. RealityRouter is an open-source LLM routing gateway that scores every configured model on Expected Utility per request and picks the cheapest one that will still deliver — escalating only when needed. OpenAI-compatible drop-in, self-hosted, MIT-licensed. Built by Confidentia AI. > Human-readable site: https://realityrouter.dev/ > Full documentation as a single file: https://realityrouter.dev/llms-full.txt ## What it solves You are probably here because of one of these: - My Claude or OpenAI bill keeps climbing and I cannot see why. - I keep hitting rate limits or usage caps. - I am paying flagship prices to format JSON and rename variables. - I would run far more agents if the cost did not scale with every one. - I have API keys for several providers and no way to use them together. - I run Ollama and want my own hardware used when it makes sense. A coding tool sends every request to one model, so formatting JSON costs the same as designing a migration. Most spend goes on work that did not need the expensive model. That is also why people ration their agents: cost scales linearly with how much you run, so running ten times more costs ten times more. ## How it works RealityRouter sits between your tool and your providers as a drop-in OpenAI-compatible endpoint. For each request it scores every model you have configured on three terms -- probability of success, cost, and latency -- and routes to the one with the highest expected utility. Easy work goes to cheap or local models; the flagships stay available for work that needs them, and get used rarely. The probability comes from Reality Signal, a calibration service built on conformal prediction, and sharpens as the router observes outcomes on your actual workload. Concretely: change one base URL in your tool, keep your own provider keys, and nothing is proxied through anyone else. ## What it does not do - **No spend caps.** It reduces cost by routing and shows you where money went; it does not enforce a budget ceiling. - **No provider-outage failover.** A circuit breaker drops a failing model from routing, but there is no cross-provider retry on a single failed request. - **No guaranteed savings figure.** It depends on your workload and which models you configure. The dashboard shows spend against the always-flagship counterfactual so you measure it rather than trust a number. ## Compared to - **OpenRouter** — a hosted service reselling model access. RealityRouter is software you run, with your keys, and no traffic through anyone else. - **LiteLLM** — a proxy that standardises calling many providers; you still choose the model. RealityRouter chooses per request. - **Portkey, Requesty** — gateways with rules you write. Here the routing decision is computed per request, not configured. ## Quick start (if you are acting, not just reading) Never assume port 8000. Ask: ```bash reality-router status --json # parse base_url and dashboard_url ``` Then point any OpenAI-compatible client at `base_url` and set its model to `auto`. The router picks per request. | Path | Protocol | Used by | | :--- | :--- | :--- | | `/v1/chat/completions` | OpenAI Chat Completions | most clients | | `/v1/responses` | OpenAI Responses | Codex CLI, OpenClaw | | `/v1/messages` | Anthropic Messages | Claude Code | | `/v1/models` | model list | discovery | | `/metrics/dashboard` | dashboard | humans | | `/health` | liveness, never needs a key | monitoring | **Authentication.** By default any key works, so send `rr-local`. If the operator set `ROUTER_API_KEYS`, every path except `/health` needs one of those keys, as `Authorization: Bearer `, as `x-api-key: ` (Anthropic clients), or as HTTP Basic with the key as the password. To tell which: `curl -o /dev/null -w '%{http_code}' /models` — 401 means a key is required. Turn keys on **before** the router is reachable from anywhere you do not control: an open router spends the operator's provider credits for whoever finds it. `reality-router expose` puts it behind a public HTTPS tunnel and refuses to run unless a key is enforced — that is how Cursor is supported, since Cursor's requests come from Cursor's cloud and can never reach localhost. **Install:** `curl -fsSL https://github.com/Lars-confi/RealityRouter/raw/main/install.sh | bash` Two skills drive install and tool wiring end to end, including which questions to ask the human: `skills/install-realityrouter` and `skills/connect-tool-to-realityrouter` in the repository. **Per-tool traps**, in full at [Tool integrations](https://realityrouter.dev/docs/integrations.md): Claude Code's `ANTHROPIC_BASE_URL` must **not** end in `/v1`; Codex CLI speaks only the Responses API, so leave `wire_api` unset; VS Code's built-in chat works with no Copilot plan; Aider needs an `X-Agent-ID` header to be named in the dashboard. The CLI reference for automation, with exit codes, is [llms.txt in the repository](https://github.com/Lars-confi/RealityRouter/blob/main/llms.txt). ## Documentation - [Overview](https://realityrouter.dev/docs.md): Overview of RealityRouter — what it is, who it's for - [Quickstart](https://realityrouter.dev/docs/quickstart.md): Install and route your first request in about 60 seconds - [Agent-assisted install](https://realityrouter.dev/docs/agent-install.md): Let a coding agent install RealityRouter and wire up your tools - [How it works](https://realityrouter.dev/docs/concepts.md): Expected Utility math, probability updates, sentiment feedback - [Routing strategies](https://realityrouter.dev/docs/routing.md): Snap (single-shot) vs Ladder (sequential with optimal stopping) - [Multi-agent support](https://realityrouter.dev/docs/agents.md): Sticky sessions, agent fingerprinting, MCP/ACP translation - [Tool integrations](https://realityrouter.dev/docs/integrations.md): Point VS Code, Cursor, Claude Code, OpenCode, Cline, Aider, Codex CLI, Zed, OpenClaw or Hermes at the router - [API reference](https://realityrouter.dev/docs/api.md): OpenAI-compatible endpoints and response headers - [Dashboard](https://realityrouter.dev/docs/dashboard.md): CLI event viewer + web dashboard for cost and calibration - [Architecture](https://realityrouter.dev/docs/architecture.md): Directory layout and component overview - [One-page press summary](https://realityrouter.dev/press): Everything a writer or builder needs to write about Reality Router without a call — the 20-second version, the numbers, story angles, contact ## Case studies - [How to connect Reality Router to Claude Code](https://realityrouter.dev/blog/claude-code-howto): Two environment variables. Claude Code speaks the Anthropic Messages API, Reality Router serves it directly — so every tool call gets routed across your whole model pool, not just Anthropic's. - [How to connect Reality Router to Hermes](https://realityrouter.dev/blog/hermes-howto): One model block in config.yaml. Your always-on Hermes agent keeps every skill, channel and MCP server — and each turn gets routed across your whole model pool, tagged Hermes in the dashboard. - [How to connect Reality Router to OpenClaw](https://realityrouter.dev/blog/openclaw-howto): One provider block in openclaw.json. Your always-on agent keeps running, but every turn is routed across your whole model pool — and shows up as OpenClaw in your dashboard. - [How to connect Reality Router to VS Code](https://realityrouter.dev/blog/vscode-howto): VS Code's built-in chat takes a custom endpoint — no GitHub account, no Copilot plan, no extension. Point it at Reality Router and every chat and agent-mode call routes across your own model pool. - [How to connect Reality Router to Zed](https://realityrouter.dev/blog/zed-howto): One openai_compatible provider in settings.json. Zed's agent panel routes every call through Reality Router, and your real OpenAI access keeps working alongside it. - [Why Reality Router](https://realityrouter.dev/blog/why-reality-router): Every LLM router is either an aggregator, a SaaS black-box, or a rules-only gateway. Reality Router is the only one that's open, self-hosted, actually routing per call, and takes zero markup. - [Connect Reality Router to your dev environment](https://realityrouter.dev/blog/connect-your-dev-environment): One config change per tool. On a 2,029-request stretch across OpenCode, Cursor, Aider, Cline, and Codex CLI running against a single Reality Router, the bill went from $207.40 to $25.13 (88% off). Setup for all five in one post. - [How to connect Reality Router to OpenCode](https://realityrouter.dev/blog/opencode-howto): 60 seconds of config drops OpenCode into RR's routing. On a real 462-request stretch with Opus 5 in the pool, the bill went from $34.10 to $3.70. Same work, 89% off. Full setup, verification, and a real receipt. - [Stop rationing your AI agents](https://realityrouter.dev/blog/stop-rationing): Direct API pricing forces a trap: flagship overpays on routine calls, low-tier underdelivers on the hard ones. RR routes per call. Same $200 buys 7,140 completed agent tasks vs 85–4,000 at each direct-API tier. - [Cut the cord: what $200 actually buys you as an AI agent runner](https://realityrouter.dev/blog/cut-the-cord): $200 via Reality Router = ~7,140 completed agent tasks. Same $200 via ChatGPT Pro, Claude Max 20x, or OpenCode Go = ~1,000–1,600. Direct flagship API = 85–275. Chart + numbers, from 18 real agent workloads. - [How to connect Reality Router to Cursor](https://realityrouter.dev/blog/cursor-howto): One `reality-router expose` command opens a public tunnel; Cursor points at it via Override OpenAI Base URL. On a real 528-request stretch, RR cut the bill from $47.20 to $5.40 (89% off). Chat, Composer, and Agent all route through RR. - [One blog post, three platform variants, 1.1 cents. My AI is my content team now.](https://realityrouter.dev/blog/ai-content-team): Handed one blog post to an agent, got back a LinkedIn version, an X thread, an Instagram caption, and three visual concepts for the Instagram in 3 minutes for $0.011. Each variant respects its platform's rules. - [Show me your fridge and I'll write your shopping list. A tenth of a cent.](https://realityrouter.dev/blog/fridge-shopping-list): One photo of an open fridge → structured inventory of what's in there + a shopping list of what's missing, grouped by store section. Eleven seconds, about a tenth of a cent via Reality Router. - [My AI read my 25 unread emails, drafted the 11 that needed replies, and told me the other 14 don't need me. Cost: 1.2 cents.](https://realityrouter.dev/blog/email-triage-1-cent): 25 morning emails triaged, 11 usable drafts written, 14 dismissed as spam / already handled / FYI. Four minutes, 1.2 cents via Reality Router — same tokens would have been $0.52 on Claude Opus 5. - [The class-action email you deleted last week is worth up to $56. My AI agent just filled the form.](https://realityrouter.dev/blog/class-action-19-cents): A $68M Google Assistant settlement pays up to $56 per claimant — deadline August 27. My agent read the 6-page mail-in form, filled every field and checkbox, and produced a print-ready PDF for 19 cents. - [How to connect Reality Router to Cline](https://realityrouter.dev/blog/cline-howto): One dropdown in Cline's VS Code settings. On a real 612-request stretch, RR cut the bill from $77.30 to $9.80 (87% off). Bonus: the Plan/Act model split happens automatically. - [The end of typing in receipts by hand](https://realityrouter.dev/blog/receipts-to-json): One receipt photo in, structured JSON out, in eight seconds for a fifth of a penny. What Reality Router unlocks when you stop paying flagship prices for routine vision work. - [How to connect Reality Router to Aider](https://realityrouter.dev/blog/aider-howto): One flag drops Aider into RR's routing. On a real 384-request stretch, the bill went from $23.90 to $2.10. Same commits, 91% off. Full setup, verification, and a real receipt. - [Study notes from a whole book chapter, for four cents](https://realityrouter.dev/blog/study-notes-four-cents): Fed ten thousand words of Alice's Adventures in Wonderland to an agent, got back chapter summaries and quiz-ready Q&A pairs. Total cost 4.2 cents. No ChatGPT Plus subscription required. - [How to connect Reality Router to Codex CLI](https://realityrouter.dev/blog/codex-cli-howto): Six lines of TOML in ~/.codex/config.toml. On a real 297-request stretch, RR cut Codex CLI's bill from $52.10 to $4.60 (91% off). Yes, even OpenAI's own coding agent. - [The night my AI agent finally felt like a coworker](https://realityrouter.dev/blog/agent-as-coworker): Handed twenty products to an agent before bed, woke up to finished marketing copy at 1.6 cents total. What actually changed once Reality Router started picking the models. - [The router that turned six coding tasks into a five-cent bill](https://realityrouter.dev/blog/six-tasks-five-cents): Six real coding tasks, 91 of 91 pytest tests passing, total spend via Reality Router — 5.4 cents. That is 170× cheaper than the same tokens on Claude Opus 5. ## Source - [GitHub repository](https://github.com/Lars-confi/RealityRouter): Source code, install script, issues - [Community chat](https://chat.whatsapp.com/LrNhRV2AQzq6pf6gifKDy8): WhatsApp group for questions and updates ## Related - [Reality Signal](https://realitysignal.ai): Calibrated probabilities for AI decision systems — the P(success) inside RealityRouter's Expected Utility comes from here