Get started

RealityRouter

The Intelligent Decision Engine for AI Agents.

RealityRouter is a high-performance LLM routing gateway designed for the agentic era. Standard proxies pass requests through; RealityRouter uses Expected Utility Theory and real-time calibration to choose the best model for every individual prompt — balancing accuracy, cost, and latency.

In a world where model performance fluctuates and costs vary by orders of magnitude, RealityRouter puts the choice back in the hands of the user.


Why RealityRouter?

For developers building AI agents (Zed, Cursor, Claude Code, Roo Code, OpenClaw, AutoGPT), picking a model is usually a trade-off between "too expensive" (flagship models on every query) and "too unreliable" (small local models that miss). RealityRouter solves this by acting as smart middleware that:

  1. Evaluates intelligence — uses the Reality Signal™ API to estimate the probability of success for each model on the specific task at hand.
  2. Calculates utility — applies a mathematical formula to balance accuracy, cost, and speed.
  3. Enforces quality — validates tool calls before they reach your agent, catching malformed JSON, ghost tool calls, and protocol leaks.

Is this for you?

RealityRouter is probably what you want if you are saying any of these:

  • My Claude / OpenAI / Cursor bill keeps climbing and I don't know why.
  • I keep hitting rate limits or usage caps.
  • I'm paying flagship prices to format JSON.
  • I'd run far more agents if the cost didn't scale with every one.
  • I can't leave an agent running overnight without watching the meter.
  • I have keys for several providers and no way to use them together.
  • I run Ollama and want to use my own hardware when it makes sense.
  • I can't tell which model is actually earning its cost.
  • I want a router I host myself, with my keys, not a service in the middle.

It is not a way to cap your spend — there is no budget ceiling — and it will not tell you which single model to standardise on. It removes that decision instead, per request.

FAQ · Quickstart · Agent-assisted install


The core engine — Expected Utility

Every request is passed through a decision-theoretic engine. For each candidate model m in your configured pool, RealityRouter computes:

EU(m) = p · R − α · cost − β · latency
  • p — calibrated probability of success on this specific prompt.
  • R — constant reward for a correct answer.
  • α — your cost sensitivity (tune in the dashboard).
  • β — your latency sensitivity (tune in the dashboard).

The router picks argmax EU — every model, every query.

Dynamic tuning. The dashboard exposes α and β as live sliders. Slide left for maximum frugality; slide right for raw speed. Changes take effect on the next request — no restart, no redeploy.


Validation gateway

RealityRouter sits between the model and your agent to enforce protocol compliance:

  • Leak protection — detects and scrubs raw tool tags that models accidentally leak into text output.
  • Ghost tool detection — rejects responses that call tools you didn't expose to the model.
  • Heuristic rescue — recovers valid JSON tool calls buried in conversational fluff.
  • Schema validation — validates tool arguments against your schema before the agent ever sees them.

Key features

  • Two routing strategies — single-shot (route to argmax-EU model immediately) or sequential (start cheap, escalate on validation failure).
  • Automatic feedback loop — validated outcomes feed back into the calibration engine, sharpening future routing decisions.
  • Multi-provider auto-discovery — bring your own keys; the router discovers and benchmarks models from OpenAI, Anthropic, Gemini, Mistral, DeepSeek, Moonshot (Kimi), Z.ai (GLM), xAI (Grok), Alibaba Qwen, local Ollama, and any OpenAI-compatible endpoint.
  • Live dashboard — track unit economics, savings vs. always-flagship, per-model reliability, and per-agent activity in a built-in web UI.

Documentation

  • Quickstart — install and route your first request in about 60 seconds.
  • Architecture — directory layout and component overview.
  • How it works — Expected Utility math and the calibration feedback loop.
  • Routing strategies — single-shot vs. sequential.
  • Multi-agent support — protocol detection and sticky sessions.
  • Agent-assisted install — let a coding agent install and wire it up for you (Claude Code: /plugin marketplace add Lars-confi/RealityRouter then /plugin install reality-router@confidentia).
  • API reference — OpenAI-compatible endpoints.
  • Dashboard — CLI event viewer + web dashboard.
  • FAQ — cost, model choice, privacy, and how this compares to other routers.

Developer integration

RealityRouter exposes an OpenAI-compatible API. The supported endpoints and features are listed in the API reference. You don't need to rewrite your agent — just change your environment variables:

  • Base URL — http://localhost:8000/v1 (obtain dynamically via reality-router status)
  • API Key — any non-empty placeholder value for local ingress in the current version. RealityRouter ingress does not currently validate this value; see SECURITY.md before remote deployment.
  • Model — auto (to let the router choose the best model) or a specific model ID (which pins the request to that model, bypassing initial selection while still supporting validation and escalation)

For AI Agents

RealityRouter is fully agent-native and designed to be managed and controlled programmatically.

If you are an AI coding agent (like Zed, Claude Code, Cursor, Roo Code, etc.) or an automation script, you can interact with RealityRouter headlessly:

  1. Zero-Prompt Setup & Autodetection:
    • Run reality-router setup --agent to configure RealityRouter. This will automatically scan your environment variables (OPENAI_API_KEY, GEMINI_API_KEY, etc.), autodetect local Ollama, choose default coefficient penalties, and complete the installation without any prompts.
  2. Strict Validation Mode:
    • Run reality-router setup --non-interactive to assert that all configurations are present. If API credentials or models are missing, it will crash with documented, non-zero exit codes.
  3. SSO Device-Code Auth over JSON Streams:
    • Run reality-router auth --agent --json to start OAuth authentication. It prints newline-delimited JSON (NDJSON) events (e.g. auth_required with verification_uri and user_code) and never logs sensitive tokens.
  4. Active Diagnostics Check:
    • Run reality-router doctor --json to perform active checks. The command exits with precise exit codes (0 if healthy, 10 if config file missing, 11 if auth missing, 12 if provider keys missing, 14 if port is busy).
  5. Machine-Readable Status:
    • Run reality-router status --json to get detailed stats about the daemon status, port, pid, and active provider models.

License

RealityRouter is MIT licensed. See the LICENSE file for the full text.


Contributing

We're building user-centric AI infrastructure. If you're interested in decision theory, agent protocols, or high-performance routing, contributions are welcome.

Built by Confidentia AI and the open-source community.