The right model for every prompt.
Automatically.
Point any agent at RealityRouter and cut your agent bill by up to 80%. Easy requests go to cheap models, hard ones go to the flagships — and you always get the right model for the job, even when you don't know which one that is.
Works with VS Code · Cursor · Claude Code · Zed · OpenCode · Aider · Cline · Codex CLI · Hermes · your own scripts
Routes to OpenAI · Anthropic · Gemini · DeepSeek · Ollama · any OpenAI-compatible endpoint
Without a router
You default to Opus or GPT flagship models for trivial tasks. 20× the cost, same answer.
Even Claude Code hits the wall mid-task. Agents stall. You wait for the window to reset.
When Anthropic lags or OpenAI 500s, your whole stack follows. No fallback, no provider competition.
With RealityRouter
One drop-in proxy between you and every LLM.
Every prompt from every agent flows through RealityRouter. It scores every model on Expected Utility — probability of success, cost, latency — and picks the most rewarding one for that specific query.
from Cursor · Claude Code · Zed · OpenCode · Aider · Cline · Codex CLI
Easy prompts go to cheap models. Hard ones go to flagships. Every prompt scored against every model, with calibrated probabilities from Reality Signal™ — automatically.
Where the probability comes from
Reality Signal™ is built on Venn predictors —
and their inventor sits on our scientific board.
Vladimir Vovk, Royal Holloway, University of London.
Vovk and Prof. Alexander Gammerman, co-founders of conformal prediction at Royal Holloway, sit on Confidentia's scientific board.
“We call them Venn predictors, and so the advantage is that you get perfect calibration for free.”
“But if you want this kind of guarantees, if you want the property of being perfectly calibrated, Venn predictors are the only way to achieve it.”
See where your money goes
And where you save it.
Every routing decision is logged with its full utility breakdown. The web dashboard shows you cost vs. counterfactual, per-model reliability, per-agent spend, and live calibration health.
Real numbers from a real install: $147 saved against the always-flagship counterfactual on 1,097 requests across Zed, RooCode, and a Python client. 95% success rate.
What a routing split looks like
| Model | Calls | Share | Cost |
|---|---|---|---|
| qwen · localqwen3-coder:30b · local | 220 | 20% | $0.00 |
| gemini-flashgemini-3.8-flash | 400 | 36% | $14.80 |
| DeepSeek FlashDeepSeek v4 Flash | 285 | 26% | $5.80 |
| gemini-progemini-3.8-pro | 140 | 13% | $32.30 |
| Opus 5Claude Opus 5 | 52 | 5% | $18.91 |
| Total | 1,097 | 100% | $71.81 |
83% of calls went to models costing under $0.04 each, and consumed 29% of the spend. The expensive models stayed available and took the other 17%.
The same 1,097 calls, if every one had gone to a single top model
| Claude Sonnet 5 | $160 | 2.2× |
| GPT-5.6 | $319 | 4.4× |
| Claude Opus 5 | $399 | 5.6× |
See it route
Watch how RealityRouter picks the right model for each kind of query — and gets sharper over time.
Cheap when it's enough
Your routine queries don't need a $10 model.
Most agentic queries are routine — formatting, small refactors, quick lookups. The router scores them in milliseconds and ships them to whatever's cheapest that won't fail. Self hosted models often win, costing you nothing.
Smart when it counts
When you really need heavy lifting, you get it.
The math isn't just "pick cheapest." It's argmax(P × Reward − α × Cost − β × Time). When a query will likely fail on a small model, P(success) collapses for cheap options and Opus wins despite the cost — and you don't lose hours debugging a wrong answer.
Self-correcting
Broken answers never reach your agent.
RealityRouter always picks the model with highest expected utility, validates the output for protocol issues — truncated JSON, AI refusals, mid-word cuts — and silently escalates if the output won't survive contact with your agent. Your client sees only the good response.
RealityRouter learns
Your frustration becomes a training signal.
When you correct a model in the next turn, RealityRouter reads it as negative feedback and lowers that model's P(success) for similar tasks. Tomorrow's routing reflects yesterday's complaints, not only yours but also your peers — without you ever filling out a survey.
The math, not the magic
Expected Utility, every query.
The router picks argmax EU — every model, every query. P(success) comes from Reality Signal™, calibrated against historical outcomes for similar tasks. α, β are your sensitivities — tune them according to your preferences, the router stays loyal forever.
Built in the open
Self-hosted. Auditable. Yours.
RealityRouter runs on your laptop, your server, or your VPC. Your API keys live in your .env. Your logs live on disk. No telemetry, no vendor lock-in.
- ✓CursorOverride OpenAI Base URL — needs a public router address · guide →
- ✓Claude CodeAnthropic-compatible passthrough · guide →
- ✓VS CodeCustom Endpoint in built-in chat, no Copilot plan · guide →
- ✓OpenCodeProvider block in opencode.json · guide →
- ✓ClineOpenAI Compatible provider in settings · guide →
- ✓Aider--openai-api-base flag or .aider.conf.yml · guide →
- ✓Codex CLImodel_providers block in config.toml · guide →
- ✓Zedopenai_compatible provider in settings.json · guide →
- ✓OpenClawProvider block in openclaw.json · guide →
- ✓Hermesmodel block in config.yaml · guide →
- ✓Continue / VSCodiumNative session tracking
- ✓Your own scriptsAny OpenAI-compatible client
New to this? Start with the setup overview →
Ready in 60 seconds.
One install command. One URL change. Every agent gets smarter.
Open · Self-hosted · Powered by Reality Signal™