Measured, not estimated

What it actually costs

Every model sat the identical exam and we counted every token it used. This is the real bill for one full pass — same questions, same graders, same day. Nobody's estimate, nobody's marketing page.

The bill

One exam, difference

Shorter bar is cheaper. Every figure here is calculated from token counts we measured, at list prices we checked ourselves — and you can put your own prices in at the bottom.

Cost of one full run

Hover any bar for the token breakdown.
Anthropic OpenAI Self-hosted

The bit that matters

What you get for the money

Score against cost, every setup. The best value sits on the left edge — nothing cheaper scores higher. Everything up and to the right of that line is paying more for the same marks or worse.

Score against cost

Anthropic OpenAI Self-hosted

One receipt worth reading

The same model, asked the same question, twice

Most of these findings need a chart. This one needs a table and a deep breath.

What actually happened
Make it your numbers

Don't like our prices? Change them.

Every dollar on this site is token counts we measured multiplied by the list prices below. Edit any of them and the whole page recalculates on the spot — handy if you're on a discount, a different region, or a vendor has moved a price since we last looked.

Model $ / M input $ / M output $ / M cached Source

⚡ And the ones with no price at all

Two of the setups on this page run on a desktop PC in somebody's office. There is no price per million tokens, because there is no invoice — there's a power bill. This model eats kilowatts, not tokens. So we costed the power bill instead, and you can argue with both halves of it:

The wattage is a measured GPU reading plus component estimates for the rest of the machine — a real number for the card, an honest guess for everything around it. It excludes the hardware, its depreciation, the room it sits in, and the hour and a half you spend waiting. If you want the true cost of self-hosting, that list is where it hides.

🧮
How the sum works. Cost = input tokens × input rate + output tokens × output rate + cached-read tokens × cached rate + cache-write tokens × input rate × .

Cache traffic is most of the bill on this workload, not a rounding error — a benchmark re-reads a big cached prompt on every single run, and ignoring it under-counted one model by 67%. The cache-write multiplier is the vendor's published 1-hour rate, which we also recovered independently at 2.002× by fitting our own spend.

Checked against Anthropic's own reported spend, this arithmetic lands within %. We publish that gap rather than tuning it away. OpenAI doesn't report a per-run cost at all, so those figures are ours alone — same method, no vendor number to check against.