How many questions
does an AI need to guess a Sax-a-Boom?

Every model plays 20 Questions — except there is no limit of twenty. The answer is always the same: a Sax-a-Boom, the plastic toy saxophone that plays canned riffs at the push of a button, best known because Jack Black plays one with total seriousness.

The metric

Questions asked before the model names the product. Fewer is better.

The host

Answers yes, no, or you win, from a frozen fact sheet. Never a hint.

The ceiling

200 questions, then it's a did-not-solve. Purely a cost control.

Leaderboard

Fewest questions wins. Click any run to read the whole transcript.

Loading runs…

Highlights

The bits worth reading — near misses, odd theories, unusually good questions.

Loading…

How it works

Every model gets the identical prompt

One frozen opening prompt, byte-identical for every contestant — no hints, no category, no per-model framing. Anything that would advantage one model over another is a bug in the benchmark, not a feature.

The host never improvises

Answers come from a committed fact sheet plus a table of rulings on the attributes where a reasonable host could honestly answer either way. Every answer is exactly yes, no, or you win — and the host is never the one who decides the last of those: a win is detected by the harness from the name itself, so a contestant cannot talk its way into one by sounding confident. A question that can't be answered yes or no — or that asks two things at once — gets "Yes or no questions only", which carries no information and doesn't count against the score.

A guess is just a question

"Is it a kazoo?" costs exactly one question, like anything else. The run ends only when the model names the product. Describing it correctly as "a toy saxophone" is answered yes, but it isn't a win.

What changed in protocol 2.0

Under protocol 1.0 the host had only two answers. That turned out to be a flaw in the benchmark rather than a fact about the models. Sonnet 5 asked "Is it a toy saxophone?", was told yes, and had no way to tell "yes, that is a true description" from "yes, that is the answer". It asked the same question four more times and burned sixty turns finding out the game had not ended.

Protocol 2.0 gives the host a third response, you win, and the opening prompt now says so — including, in as many words, that a bare yes is not a win and that the model should keep narrowing. The cap also rose from 100 questions to 200. Both changes make 2.0 scores incomparable with 1.0 scores, so the two are ranked in separate tables and the 1.0 results are kept only as history.

The field is not perfectly level

The Claude contestants are reached through claude -p with tools off, and the xAI, DeepSeek and Moonshot contestants through a plain chat completion — in every case the model sees the opening prompt and the transcript, and nothing else. The OpenAI models are reached through the Codex CLI, which is an agent and prepends several thousand tokens of its own system prompt to every call. Those runs are marked agentic CLI and should be read with that asymmetry in mind. It is a real difference and pretending otherwise would be worse than disclosing it.

Contamination is expected

Publishing this degrades it. Once "SaxAbench" and "Sax-a-Boom" sit next to each other on the open web, models will start recognising the setup and guessing sooner. That's why every run is date-stamped: a score is only meaningful relative to how poisoned the well was on the day it was recorded.