Get started
RealityRouter
The Intelligent Decision Engine for AI Agents.
RealityRouter is a high-performance LLM routing gateway designed for the agentic era. Standard proxies pass requests through; RealityRouter uses Expected Utility Theory and real-time calibration to choose the best model for every individual prompt — balancing accuracy, cost, and latency.
In a world where model performance fluctuates and costs vary by orders of magnitude, RealityRouter puts the choice back in the hands of the user.
Why RealityRouter?
For developers building AI agents (Zed, Cursor, Claude Code, Roo Code, OpenClaw, AutoGPT), picking a model is usually a trade-off between "too expensive" (flagship models on every query) and "too unreliable" (small local models that miss). RealityRouter solves this by acting as smart middleware that:
- Evaluates intelligence — uses the Reality Signal™ API to estimate the probability of success for each model on the specific task at hand.
- Calculates utility — applies a mathematical formula to balance accuracy, cost, and speed.
- Enforces quality — validates tool calls before they reach your agent, catching malformed JSON, ghost tool calls, and protocol leaks.
Is this for you?
RealityRouter is probably what you want if you are saying any of these:
- My Claude / OpenAI / Cursor bill keeps climbing and I don't know why.
- I keep hitting rate limits or usage caps.
- I'm paying flagship prices to format JSON.
- I'd run far more agents if the cost didn't scale with every one.
- I can't leave an agent running overnight without watching the meter.
- I have keys for several providers and no way to use them together.
- I run Ollama and want to use my own hardware when it makes sense.
- I can't tell which model is actually earning its cost.
- I want a router I host myself, with my keys, not a service in the middle.
It is not a way to cap your spend — there is no budget ceiling — and it will not tell you which single model to standardise on. It removes that decision instead, per request.
FAQ · Quickstart · Agent-assisted install
The core engine — Expected Utility
Every request is passed through a decision-theoretic engine. For each
candidate model m in your configured pool, RealityRouter computes:
EU(m) = p · R − α · cost − β · latency
p— calibrated probability of success on this specific prompt.R— constant reward for a correct answer.α— your cost sensitivity (tune in the dashboard).β— your latency sensitivity (tune in the dashboard).
The router picks argmax EU — every model, every query.
Dynamic tuning. The dashboard exposes
αandβas live sliders. Slide left for maximum frugality; slide right for raw speed. Changes take effect on the next request — no restart, no redeploy.
Validation gateway
RealityRouter sits between the model and your agent to enforce protocol compliance:
- Leak protection — detects and scrubs raw tool tags that models accidentally leak into text output.
- Ghost tool detection — rejects responses that call tools you didn't expose to the model.
- Heuristic rescue — recovers valid JSON tool calls buried in conversational fluff.
- Schema validation — validates tool arguments against your schema before the agent ever sees them.
Key features
- Two routing strategies — single-shot (route to argmax-EU model immediately) or sequential (start cheap, escalate on validation failure).
- Automatic feedback loop — validated outcomes feed back into the calibration engine, sharpening future routing decisions.
- Multi-provider auto-discovery — bring your own keys; the router discovers and benchmarks models from OpenAI, Anthropic, Gemini, Mistral, DeepSeek, Moonshot (Kimi), Z.ai (GLM), xAI (Grok), Alibaba Qwen, local Ollama, and any OpenAI-compatible endpoint.
- Live dashboard — track unit economics, savings vs. always-flagship, per-model reliability, and per-agent activity in a built-in web UI.
Documentation
- Quickstart — install and route your first request in about 60 seconds.
- Architecture — directory layout and component overview.
- How it works — Expected Utility math and the calibration feedback loop.
- Routing strategies — single-shot vs. sequential.
- Multi-agent support — protocol detection and sticky sessions.
- Agent-assisted install — let a coding agent
install and wire it up for you (Claude Code:
/plugin marketplace add Lars-confi/RealityRouterthen/plugin install reality-router@confidentia). - API reference — OpenAI-compatible endpoints.
- Dashboard — CLI event viewer + web dashboard.
- FAQ — cost, model choice, privacy, and how this compares to other routers.
Developer integration
RealityRouter exposes an OpenAI-compatible API. The supported endpoints and features are listed in the API reference. You don't need to rewrite your agent — just change your environment variables:
- Base URL —
http://localhost:8000/v1(obtain dynamically viareality-router status) - API Key — any non-empty placeholder value for local ingress in the current version. RealityRouter ingress does not currently validate this value; see
SECURITY.mdbefore remote deployment. - Model —
auto(to let the router choose the best model) or a specific model ID (which pins the request to that model, bypassing initial selection while still supporting validation and escalation)
For AI Agents
RealityRouter is fully agent-native and designed to be managed and controlled programmatically.
If you are an AI coding agent (like Zed, Claude Code, Cursor, Roo Code, etc.) or an automation script, you can interact with RealityRouter headlessly:
- Zero-Prompt Setup & Autodetection:
- Run
reality-router setup --agentto configure RealityRouter. This will automatically scan your environment variables (OPENAI_API_KEY,GEMINI_API_KEY, etc.), autodetect local Ollama, choose default coefficient penalties, and complete the installation without any prompts.
- Run
- Strict Validation Mode:
- Run
reality-router setup --non-interactiveto assert that all configurations are present. If API credentials or models are missing, it will crash with documented, non-zero exit codes.
- Run
- SSO Device-Code Auth over JSON Streams:
- Run
reality-router auth --agent --jsonto start OAuth authentication. It prints newline-delimited JSON (NDJSON) events (e.g.auth_requiredwithverification_urianduser_code) and never logs sensitive tokens.
- Run
- Active Diagnostics Check:
- Run
reality-router doctor --jsonto perform active checks. The command exits with precise exit codes (0if healthy,10if config file missing,11if auth missing,12if provider keys missing,14if port is busy).
- Run
- Machine-Readable Status:
- Run
reality-router status --jsonto get detailed stats about the daemon status, port, pid, and active provider models.
- Run
License
RealityRouter is MIT licensed. See the LICENSE file for the full text.
Contributing
We're building user-centric AI infrastructure. If you're interested in decision theory, agent protocols, or high-performance routing, contributions are welcome.
Built by Confidentia AI and the open-source community.