Live Findings
What we measured, not what we estimated
Same questions, same open-source codebase, same model. One agent with Nebulai, one without. The numbers below are from the latest complete paired run.
Session cost29%
lower ($3.91 → $2.77)
Tokens48%
fewer (2,976,053 → 1,541,139)
Grader passes18/18
Nebulai · 18/18 unaided
By question
Cost saved on each question, means over three paired attempts. Both sides passed the grader on every attempt. One question cost slightly more.
| Question | Saved |
|---|---|
| Error propagation | −3% |
| Middleware composition | +27% |
| Smart router selection | +48% |
| Request parameters | +52% |
| RegExp route compilation | +18% |
| Context and response | +23% |
How to read this
- Attempts were paired and interleaved: the same question, the same attempt number, the same model, with and without Nebulai.
- Cost is the agent's own session accounting, including cache. It does not include building or hosting Project Memory.
- The grader is a mechanical checklist of required facts. It is not an independent review of whether the answer was right.
- Results vary by question. The headline is a workload total, not a guarantee on the next task.
An earlier published run on the same six questions is on Measured (4 September 2026): 21–28% lower measured-session cost and 13–24% fewer tokens on the same six Hono questions, across two repeated Claude Code runs.