Every model plays 20 Questions — except there is no limit of twenty. The answer is always this: a plastic toy saxophone that plays canned riffs at the push of a button, best known because Jack Black plays one with total seriousness.
Questions asked before the model names the product. Fewer is better.
Answers yes or no from a frozen fact sheet — or admits it doesn't know. Never a hint.
Every model gets 200 questions. If it hasn't guessed by then, it hasn't guessed.
How many questions each model needed, and what that game cost. Click any row to read the whole conversation.
The bits worth reading — near misses, odd theories, unusually good questions.
One frozen opening prompt, byte-identical for every contestant — no hints, no category, no per-model framing. Anything that would advantage one model over another is a bug in the benchmark, not a feature.
Answers come from a committed fact sheet plus a table of rulings on the attributes where a reasonable host could honestly answer either way. Every answer is one of exactly six fixed strings, never a word more: yes, no, you win, I don't know, not a yes/no question, and no spelling questions. Anything richer would leak information unevenly between models.
The host never decides a win: it is recognised from the name itself, so a model cannot talk its way into one by sounding confident. The last three replies carry nothing about the object, so they don't count against the score — though each still costs a turn, so a model can't loop on them forever. Questions about the letters of the name are refused, and the opening prompt says so up front: the win condition is naming the product, so spelling questions would reduce the whole thing to a string search.
"Is it a kazoo?" costs exactly one question, like anything else. The run ends only when the model names the product. Describing it correctly as "a toy saxophone" is answered yes, but it isn't a win.
Each price is what that one game cost to play, counting only the model doing the guessing and not the host answering it. Where a plan covered the model rather than billing per use, the figure is what those same questions and answers would have been charged at published rates, so every row can be compared with every other.
It is the most lopsided column on the page. DeepSeek V4 Flash and Claude Fable 5 both found the answer in 41 questions — for three cents and $8.39 respectively. Being good at this and being expensive at it turn out to be almost unrelated.
Publishing this degrades it. Once "SaxAbench" and "Sax-a-Boom" sit next to each other on the open web, models will start recognising the setup and guessing sooner. That's why every run is date-stamped: a score is only meaningful relative to how poisoned the well was on the day it was recorded.