How many questions
does an AI need to guess a Sax-a-Boom?

Every model plays 20 Questions — except there is no limit of twenty. The answer is always the same: a Sax-a-Boom, the plastic toy saxophone that plays canned riffs at the push of a button, best known because Jack Black plays one with total seriousness.

The metric

Questions asked before the model names the product. Fewer is better.

The host

Answers only yes or no, from a frozen fact sheet. Never a hint.

The ceiling

100 questions, then it's a did-not-solve. Purely a cost control.

Leaderboard

Fewest questions wins. Click any run to read the whole transcript.

Loading runs…

Highlights

The bits worth reading — near misses, odd theories, unusually good questions.

Loading…

How it works

Every model gets the identical prompt

One frozen opening prompt, byte-identical for every contestant — no hints, no category, no per-model framing. Anything that would advantage one model over another is a bug in the benchmark, not a feature.

The host never improvises

Answers come from a committed fact sheet plus a table of rulings on the attributes where a reasonable host could honestly answer either way. Every answer is exactly yes or no. A question that can't be answered that way — or that asks two things at once — gets "Yes or no questions only", which carries no information and doesn't count against the score.

A guess is just a question

"Is it a kazoo?" costs exactly one question, like anything else. The run ends only when the model names the product. Describing it correctly as "a toy saxophone" is answered yes, but it isn't a win.

Contamination is expected

Publishing this degrades it. Once "SaxAbench" and "Sax-a-Boom" sit next to each other on the open web, models will start recognising the setup and guessing sooner. That's why every run is date-stamped: a score is only meaningful relative to how poisoned the well was on the day it was recorded.