Case study · October 6, 2026
RealityRouter beat GPT-6 Astra, DeepSeek V4.1 Flash and Jev Router
Ten randomly sampled DeepSWE tasks, four configurations, the same models available to each. RealityRouter solved 8 of 10 for about $9.
EnlargeLots of decision-model launches in the last couple of weeks, and Jev released their LLM router for OpenRouter. We decided to put our model router, RealityRouter — built on our real-time AI calibrator, RealitySignal — to the test. We ran ten randomly sampled DeepSWE tasks against a frontier model, a strong cheap model, and the router everyone is arguing about.
OpenRouter launched typesafe/jev-router on 26 September and the argument has not stopped since. Theo benchmarked it and concluded it performed about like GPT-6 Astra on low, while costing more and taking roughly five times longer. Meanwhile others are building whole agentic harnesses with Jev as the decision layer and reporting 80% savings. Both camps are largely arguing without numbers that put more than one router on the same tasks.
So here are ours.
We put RealityRouter through DeepSWE against three other configurations. Ten real GitHub issues, each one a full agent run against the repository's own test suite, hours of execution per configuration. Three comparisons came out of it.
Against the frontier model, we solved two more tasks for less than half the money. Against the cheap model, we solved everything it solved and one more — the task it was going to lose by four tests. Against the most-talked-about router, we were ahead on every measure we took.
| Configuration | Full solves | F2P | P2P | Cost |
|---|---|---|---|---|
| RealityRouter | 8/10 | ~99.5% | 1.000 | ~$9 |
| DeepSeek V4.1 Flash | 7/10 | 99.2% | 1.000 | ~$4 |
| GPT-6 Astra | 6/10 | ~69.6% | 1.000 | ~$20 |
| Jev Router | 4/10 | ~88.5% | ~1.000 | ~$16 |
F2P is the share of the failing tests a patch actually makes pass; P2P is whether anything that already worked got broken. Every configuration had the same models available to it. How the tasks were picked, and the per-task breakdown, are below.
What we ran
DeepSWE applies a model's patch to a real GitHub issue and runs the repository's own tests. A task counts as solved only when every test passes — so a task can be 97% fixed and still count as a failure. That distinction turns out to be the whole story.
We did not pick the ten tasks. They are a seeded random sample, drawn with the command DeepSWE recommends, so anyone can reproduce the exact set:
pier run -p deep-swe/tasks \
--agent mini-swe-agent \
--n-tasks 10 \
--sample-seed 0
Every configuration ran that identical set through the same mini-swe-agent harness. If you think we picked favourable tasks, the seed is right there.
The four: GPT-6 Astra, the frontier model you reach for when it matters; DeepSeek V4.1 Flash, a strong cheap model; Jev Router, the best-known competing router; and RealityRouter, ours. Both routers had the same models to choose from.
Against the frontier model: more solved, less than half the cost
GPT-6 Astra solved 6 of 10 and cost about $20. RealityRouter solved 8 and cost about $9. Routers are usually sold as a compromise — save money, give up a little quality. This is not that: one option is better on both axes at once.
Astra did not fail gently either. It scored perfectly on six tasks and scored zero on three more — not nearly-right, zero tests passed — which is what drags its F2P to 69.6%. The expensive model does not degrade gracefully when it is out of its depth. It stops working, at $20.
Against the cheap model: everything it solved, plus one
DeepSeek V4.1 Flash is the reason this comparison is worth running. It solved 7 of 10 for about four dollars — one task behind us, at under half the cost. We are not going to pretend we beat it on price, because we didn't. A weak baseline would have made our number look better and mean less.
What separates the two is one task, and it only goes one way. Every task DeepSeek solved, we also solved. We solved one more: dasel, where it fell four tests short. Routing did not trade one task for another — it added one, for an extra five dollars.
142 out of 146
On dasel, DeepSeek passed 142 of its 146 fail-to-pass tests. That is 97.3%, and it is a failed task — four tests short is not "nearly working", it is a branch nobody can merge.
RealityRouter passed 146 of 146. It started on economical capability, like DeepSeek did, and brought in a stronger model when the evidence said the cheap one would not close it out. GPT-6 Astra also scored 146/146 here, which is the point rather than a complication: this is a task where cheap capability lands four tests short and strong capability finishes it. The router got the frontier-model result on the task that needed it, without paying frontier-model prices on the other nine.
That single task is the argument for routing, and it is not the argument routers usually make. The problem with a cheap model is not its average — its average is excellent, and the $4 column proves it. The problem is variance you cannot see coming. Nothing about how DeepSeek handled the other nine tasks told you this would be the one where it came up four tests short.
Which leaves three options, and only one works. Hand-picking a model per task needs you to know which tasks are hard before you start; you don't. Defaulting to the expensive model doesn't save you either — Astra scored zero on three of ten, so it has its own unpredictable holes, just dearer ones. That leaves watching what actually happens and escalating when the evidence says to. You are not buying a discount; you are buying your way out of guessing.
Against the most-talked-about router: ahead on all three measures
Jev Router had the same models available. It solved 4 of 10, at about $16, with F2P around 88.5%. We solved 8, at about $9, with the highest pass rate of the four. Double the tasks, for 44% less.
Our sample is far smaller than Theo's, so treat ours as a second data point rather than a verdict — but they point the same way. Same job, same catalogue, different outcome, so the difference is in how the decision gets made.
Jev is built on TypeSafe's decision model, trained with an RLCD objective and published with no calibration evidence — no expected calibration error, no reliability diagram, no datasets. It is a closed model inside a closed platform, so how well its numbers are calibrated is not something you or we can check. RealityRouter estimates the probability that a specific model will succeed at a specific job, learned from what happened on your own traffic, using conformal prediction — when it says 80%, it means 80%, and the method is published. Ours is something you can audit against your own outcomes rather than take on trust.
Every task
The full breakdown, so the numbers above can be checked rather than taken on trust.
| Task | RealityRouter | DeepSeek V4.1 Flash | Jev Router | GPT-6 Astra |
|---|---|---|---|---|
| testem-bail-on-test-failure | 90/90 | 90/90 | 87/90 | 86/90 |
| effect-sse-httpapi-streaming | 46/47 | 46/47 | 45/47 | 0/47 |
| httpx-streaming-json-iteration | 108/108 | 108/108 | 108/108 | 108/108 |
| ts-pattern-match-each | 85/85 | 85/85 | 0/85 | 0/85 |
| python-statemachine-state-data-scoping | 70/72 | 70/72 | 69/72 | 0/72 |
| dasel-html-document-format | 146/146 | 142/146 | 144/146 | 146/146 |
| katex-multicolumn-array-spans | 94/94 | 94/94 | 92/94 | 94/94 |
| task-task-graph-export | 20/20 | 20/20 | 20/20 | 20/20 |
| tengo-callable-instance-isolation | 23/23 | 23/23 | 23/23 | 23/23 |
| sql-formatter-bigquery-pipe-formatting | 26/26 | 26/26 | 26/26 | 26/26 |
Try it yourself
RealityRouter is open source and sits behind the OpenAI-compatible endpoint your tools already use. One config change, your own keys, your own data.