Same-task comparison · Cost and waiting
Coding with Haiku: weigh cost and waiting together
Bito gives the same coding work to five Claude configurations. Session cost, elapsed time and implementation checks shape the choice, alongside Haiku's long-prompt price boundary.
This continues our Haiku and everyday model-cost readings, using Bito's October 9 report and official prices checked October 11. Figures are BitShovel diagrams of the author's data; we did not rerun the coding tasks.
Similar cost, different waiting
Across Bito's six main coding tasks, Haiku medium costs about $0.31 per session and takes 4.8 minutes, versus Sonnet medium's $0.64 and 3.7 minutes. Haiku high reaches $0.62 and 7.3 minutes. Its 79% versus 77% is within the author's roughly three-point run-to-run noise. Lower spending and earlier delivery are separate considerations.
| Model / Claude Code effort | Score / % of maximum | USD / session | Minutes / session |
|---|---|---|---|
| Opus 5.5 · medium (test default) | 85% | $2.71 | 8.5 |
| Sonnet 5.5 · high | 82% | $1.19 | 7.6 |
| Haiku 5.5 · high | 79% | $0.62 | 7.3 |
| Sonnet 5.5 · medium (test default) | 77% | $0.64 | 3.7 |
| Haiku 5.5 · medium (test default) | 76% | $0.31 | 4.8 |
Sources and further reading
- Bito · Coding tasks across five Claude configurationsTest authors' report · October 9Tasks from its own services, grades and session usage; Opus 5.5 is both a contestant and the judge.
Looking up a setting differs from implementing and checking
The authors separate one-shot requests, short multi-turn tasks and long sessions. Haiku high approaches Opus medium on quick code lookups, while implementation and final checks widen the gap. Our reading is to identify the required endpoint: finding a clue is useful, but delivering a working change also requires checks. A blended percentage cannot decide which outcome is sufficient.
Sources and further reading
- Bito · Coding tasks across five Claude configurationsTest authors' report · October 9Tasks from its own services, grades and session usage; Opus 5.5 is both a contestant and the judge.
A growing session can cross into a higher price tier
Official rates apply per request. At up to 100,000 prompt tokens, Haiku costs $0.10/$0.50 per million input/output tokens; above that, $0.50/$2.50. Cache reads and writes count toward prompt length. A cache hit does not remove the threshold. Assess growing sessions by material per request, call count and total cost. The table's usage-based API costs do not price a Claude subscription.
Sources and further reading
- Bito · Coding tasks across five Claude configurationsTest authors' report · October 9Tasks from its own services, grades and session usage; Opus 5.5 is both a contestant and the judge.
- Anthropic · Official API pricingOfficial conditions · Checked October 11Haiku's 100,000-prompt-token boundary includes cache reads and writes and applies per request, separately from total test-session costs.
Read the task comparison with its method
Bito uses fixed prompts from its services, a pinned Claude Code version and code-checked answer keys. Opus 5.5 grades five dimensions while also competing. A companion article describes a usual four sessions per task; the main table lacks per-cell counts, the version number, full prompts and failure logs. Medium and high differ in compute budget. This informs selection without making every result independently reproducible or comparable to our AA index and Pi pass rates.
Sources and further reading
- Bito · Coding tasks across five Claude configurationsTest authors' report · October 9Tasks from its own services, grades and session usage; Opus 5.5 is both a contestant and the judge.
- Bito · Companion new/old Haiku comparisonCompanion account of the authors' test methodDescribes the same tasks, judge and answer keys: normally four sessions per task, but two for old 4.5. The main table has no per-cell run counts.
Connect savings to the outcome you need
This continues the everyday-cost question: start with bounded lookups and explanations, then observe retries and completeness during multi-turn changes. For sustained implementation and review, judge a more expensive configuration alongside waiting, rework and human checking. Compare candidates on your own shared task and delivery standard, rather than inheriting a precise rank from one study.
Sources and further reading
- Bito · Coding tasks across five Claude configurationsTest authors' report · October 9Tasks from its own services, grades and session usage; Opus 5.5 is both a contestant and the judge.