Case study · September 20, 2026
How to connect Reality Router to Hermes
One model block in config.yaml. Your always-on Hermes agent keeps every skill, channel and MCP server — and each turn gets routed across your whole model pool, tagged Hermes in the dashboard.
One model: block in config.yaml. Hermes keeps its skills, its MCP servers and its channels — every turn just gets routed per call, with receipts.
Hermes is the agent you talk to from Telegram at 11pm, that runs cron jobs you forgot about, and that reaches a dozen MCP servers to answer a question you asked in six words. It is, in other words, the least predictable model spend you own: you cannot guess in advance whether the next turn is a one-line answer or a twenty-step tool chain.
That is the exact case for routing per call rather than picking a model up front. Point Hermes at Reality Router and:
- Each turn is scored on its own. Short answers go to small models, heavy tool chains escalate.
- Every provider you own is in reach, from one config block, instead of switching Hermes between providers by hand.
- Hermes gets its own line in the dashboard, tagged
Hermes, separate from your coding tools — so the always-on agent's cost stops hiding inside your total.
1. Point Hermes at the router (1 minute)
Edit ~/.hermes/config.yaml:
model:
provider: custom
base_url: http://localhost:8000/v1
api_key: rr-local
api_mode: chat_completions
default: auto
context_length: 128000
provider: customis what selects an OpenAI-compatible endpoint. Hermes will report its provider ascustomrather than by name — the reliable indicator of where it's pointing isbase_url.default: autohands the choice to the router per call. Any id fromreality-router modelspins it instead.api_keycan be any placeholder for a local router.
Restart the gateway if you run one, and hermes model will show the new endpoint.
2. Verify it worked (30 seconds)
hermes chat -q "Create a file called hermes.txt containing: routed via RR. Then read it back and tell me the contents."
That exercises the whole path — tool definitions out, tool call back, result returned, agent finishes. Then open the dashboard at http://localhost:8000/metrics/dashboard and look for Hermes in Agent Activity.
You'll see several requests for that one instruction. Each tool step is separately routed, which is the point.
If you run a lot of MCP servers
Worth knowing before you route a heavily-equipped Hermes: OpenAI models reject any request carrying more than 128 tool definitions. A Hermes instance with a dozen MCP servers connected can pass that on its own — each MCP adds its own tools plus four automatic ones, and Hermes's built-in toolsets add around thirty more.
The router handles it: a model that rejects the request fails and it falls through to one that accepts it (Anthropic and DeepSeek models allow considerably more). But if your pool is mostly OpenAI models, you'll be paying for a lot of failed first attempts. Two options:
- Trim the toolsets Hermes sends, per platform, with
hermes tools. - Keep models with higher tool limits in the pool so there's always somewhere to land.
This isn't a routing problem — it's the same limit you'd hit calling OpenAI directly. It's just more visible once every call is logged.
Why bother
A personal agent has a cost shape nobody plans for. It answers all day, and each answer is cheap enough to ignore, which is precisely how the bill grows without a culprit.
- The floor drops. Reminders, lookups and one-line answers stop being billed at flagship rates.
- The ceiling stays. Hard turns still reach a strong model; you just stop paying for capability you didn't need.
- You can finally see it. Per request, per model, tagged
Hermes, alongside every other tool you route — so "what does my agent actually cost" becomes a number instead of a guess.
Same agent, same skills, same channels. Routed and visible.
Reality Router is open source and self-hosted. Setup verified with Hermes Agent 0.10.0 against a running Reality Router, September 2026.