How we tested
This is the page where we hand you the stick to hit us with. Every method, every caveat, and the numbers that make our own conclusions look shakier.
"An AI graded this. Of course it picked its favorite."
Fair. So we had two AIs mark every answer — one from each company — separately, against the same written answer key, neither knowing what the other said. Then we checked whether either one went easy on its own side.
—
Half our own questions barely tell them apart
A benchmark question that everyone gets right measures nothing. Here's how well each of our questions actually separates the field — including the ones that don't, which we're leaving in and labeling rather than quietly dropping.
How many times we actually ran each thing
Much of this board rests on a single run. A chart that hides that is overclaiming, so here it is. More runs means we know how consistent a model is, which is a different question from how good it is.
The boring but important bit
If you don't say how you measured, you haven't published a result — you've published an opinion with numbers attached.
🛠️ Tools stay switched on
Both sides can search the web and use their tools, because that's how people actually use them. Testing with tools off measures a setup nobody runs. We mark the answer and record what it cost.
🔒 One question, one fresh chat
First answer counts. No retries, no re-wording, no "try again but better." That's how you'd actually use it.
🙈 Marked blind
Answers are saved under random filenames. The model's name lives only in the results file and gets joined back after marking is done.
🧾 We store tokens, not dollars
Prices move constantly. Baking them into the data would make every old result un-recalculable. Costs are applied at display time from a dated price table — which you can edit.
🖼️ Identical images
Picture questions use the same pre-rendered file for every model. Re-screenshotting per model would mean different crops and meaningless comparisons.
📖 Answer keys stay shut
Nobody reads the key until collection is finished. Knowing the answer changes how you mark an ambiguous response, and you won't notice it happening.
What you should be suspicious of
Every question, every setup
Nothing above is a vibe. Pick any two setups and compare them question by question. Turn on Nerd mode for the graders' own reasoning.
| ID | Question | Skill | A | B |
|---|
Remember when Windows 10 came out?
We were told it was "better." We were told it was "more secure." What we were never told was what was better, or which part was more secure, or how anyone would ever know. It took years for regular people to work out what had actually changed.
And we still miss XP, damn it.
AI is doing the exact same thing right now — "more capable," "more aligned," "state of the art" — and the models themselves are the worst possible source, because every one of them will cheerfully tell you it's excellent at something it just failed. So we stopped asking and started marking.