Skip to content
Live Findings

What we measured, not what we estimated

Same questions, same open-source codebase, same model. One agent with Nebulai, one without. The numbers below are from the latest complete paired run.

10 September 2026 · honojs/hono · Claude Sonnet · 6 questions · 3 repeats · 18 paired attempts

Session cost29%

lower ($3.91$2.77)

Tokens48%

fewer (2,976,0531,541,139)

Grader passes18/18

Nebulai · 18/18 unaided

By question

Cost saved on each question, means over three paired attempts. Both sides passed the grader on every attempt. One question cost slightly more.

Latest complete run · 10 September 2026 · session cost vs unaided agent
QuestionSaved
Error propagation−3%
Middleware composition+27%
Smart router selection+48%
Request parameters+52%
RegExp route compilation+18%
Context and response+23%

How to read this

  • Attempts were paired and interleaved: the same question, the same attempt number, the same model, with and without Nebulai.
  • Cost is the agent's own session accounting, including cache. It does not include building or hosting Project Memory.
  • The grader is a mechanical checklist of required facts. It is not an independent review of whether the answer was right.
  • Results vary by question. The headline is a workload total, not a guarantee on the next task.

An earlier published run on the same six questions is on Measured (4 September 2026): 21–28% lower measured-session cost and 13–24% fewer tokens on the same six Hono questions, across two repeated Claude Code runs.