Every model plays 20 Questions — except there is no limit of twenty. The answer is always the same: a Sax-a-Boom, the plastic toy saxophone that plays canned riffs at the push of a button, best known because Jack Black plays one with total seriousness.
Questions asked before the model names the product. Fewer is better.
Answers only yes or no, from a frozen fact sheet. Never a hint.
100 questions, then it's a did-not-solve. Purely a cost control.
Fewest questions wins. Click any run to read the whole transcript.
The bits worth reading — near misses, odd theories, unusually good questions.
One frozen opening prompt, byte-identical for every contestant — no hints, no category, no per-model framing. Anything that would advantage one model over another is a bug in the benchmark, not a feature.
Answers come from a committed fact sheet plus a table of rulings on the attributes where a reasonable host could honestly answer either way. Every answer is exactly yes or no. A question that can't be answered that way — or that asks two things at once — gets "Yes or no questions only", which carries no information and doesn't count against the score.
"Is it a kazoo?" costs exactly one question, like anything else. The run ends only when the model names the product. Describing it correctly as "a toy saxophone" is answered yes, but it isn't a win.
Publishing this degrades it. Once "SaxAbench" and "Sax-a-Boom" sit next to each other on the open web, models will start recognising the setup and guessing sooner. That's why every run is date-stamped: a score is only meaningful relative to how poisoned the well was on the day it was recorded.