Every model plays 20 Questions β except there is no limit of twenty. The answer is always this: a plastic toy saxophone that plays canned riffs at the push of a button, best known because Jack Black plays one with total seriousness.
Questions asked before the model names the product. Fewer is better.
Answers yes or no from a frozen fact sheet β or admits it doesn't know. Never a hint.
200 questions, then it's a did-not-solve. Purely a cost control.
How often the model solved it, then how many questions it needed when it did. Click any row to read a whole transcript.
The bits worth reading β near misses, odd theories, unusually good questions.
One frozen opening prompt, byte-identical for every contestant β no hints, no category, no per-model framing. Anything that would advantage one model over another is a bug in the benchmark, not a feature.
Answers come from a committed fact sheet plus a table of rulings on the attributes where a reasonable host could honestly answer either way. Every answer is one of exactly six fixed strings, never a word more: yes, no, you win, I don't know, not a yes/no question, and no spelling questions. Anything richer would leak information unevenly between models.
The host is never the one who decides a win: it is detected by the harness from the name itself, so a contestant cannot talk its way into one by sounding confident. The last three replies carry nothing about the object, so they don't count against the score β though each still costs a turn, so a model can't loop on them forever. Questions about the letters of the name are refused, and the opening prompt says so up front: the win condition is naming the product, so spelling questions would reduce the whole thing to a string search.
"Is it a kazoo?" costs exactly one question, like anything else. The run ends only when the model names the product. Describing it correctly as "a toy saxophone" is answered yes, but it isn't a win.
Every run now shows what the model itself cost, never what the host cost to referee it. A figure marked * is what the run would have cost on the public API β those models are covered by a subscription, so no money actually changed hands. Unmarked figures are real spend.
It is the most lopsided column on the page. DeepSeek V4 Flash and Claude Fable 5 both found the answer in 41 questions β for three cents and $8.39 respectively. Being good at this and being expensive at it turn out to be almost unrelated.
The host is a language model reading from a fixed fact sheet, and it is not perfect. Roughly one turn in seventy it declines a question it could have answered, or refuses one that was perfectly well formed. Those turns are charged to the model, which is unfair to it.
That is worth stating rather than implying away β but it is also a large improvement. Under the previous rules the host wrongly refused 140 of 291 rejected questions, roughly half of them, because it had no way to say "I can't answer that" and said "that isn't a proper question" instead. One model lost 40% of its run to it.
The host has one more answer than yes and no: it can say I don't know when the fact sheet genuinely doesn't settle a question. That is honest, and it replaced something worse β previously such questions were rejected as though they were badly formed, so models rephrased them over and over.
But it costs the model a question either way, and models did not draw them evenly. The two that never finished collected them at 14% of their turns; the joint-fastest drew none at all. One model was asked, on its fourth question, whether the thing is used indoors β a broad and useful question β and got I don't know. It went on to need 134 questions, having needed 26 under the previous rules.
So part of a score is luck about which questions happened to fall in a gap in the answer key. Treat the top of the board as a rough grouping rather than a ranking, and treat gaps of a few questions as noise.
It is tempting to blame a single unlucky answer for a run like that, so we checked. Across its three runs Claude Opus 5 took 20, 30 and 59 questions to establish it was looking at a toy saxophone, and then 38, 41 and 66 more to name it. Both halves stretched together. The slow run was not derailed at one moment β it searched less efficiently throughout, against an identical prompt and an identical answer key.
DeepSeek V4 Pro, DeepSeek V4 Flash and Kimi K2.6 all ended on runs of empty replies. That is a limitation of this harness and should not be read as any of them giving up. Every contestant reached over a plain chat API is sent the same 6,000-token output budget, and for a reasoning model that budget covers its thinking as well as its answer. Dozens of their turns came back having spent exactly 6,000 tokens β the cap β with nothing left for the question itself. They were thinking, not declining.
This is not a guess. We replayed the exact prompt DeepSeek was answering when it went blank, changing nothing but the budget:
max_tokens = 6,000 β 6,000 output tokens, empty reply max_tokens = 32,000 β 6,394 output tokens, "Is it a Chicco product?"
It needed 394 tokens more than the cap we set. Same model, same prompt, same moment in the game β the only difference between a blank and a perfectly good question was our number.
Worse, that budget is not applied evenly. The Claude and OpenAI contestants run through their own command-line tools with no output cap set by this benchmark, so they were never exposed to the failure at all. The models most penalised are the ones that reason the most. DeepSeek V4 Pro had correctly narrowed to a battery-powered plastic toy instrument by question 100 and still scored nothing.
These runs were deliberately not re-run with a larger budget. That would have handed one model more room to think than everything already scored. Raising the limit for everyone and re-running from scratch is the right fix, and that is what the current leaderboard does.
Read the leaderboard accordingly: every model that failed this way should be treated as unscored, not as beaten. Three models across two providers hit it, which makes it a defect in the protocol rather than a property of any one of them.
Claude Haiku 4.5 finished a full 200-question run under these rules. It took six sittings across separate five-hour windows, with the run pausing and resuming each time the subscription quota ran out. Its replicate runs are still working through that same process.
The reason is a setting we chose, not anything about how fast Haiku answers. At the reasoning effort we run, it bills about 4,900 output tokens per turn β roughly 34x what Opus bills β and about 99.8% of that is reasoning we asked for rather than the question it finally asks. The questions themselves are as short and as sensible as anyone else's. We are paying for thinking we never see, and then reading the bill as though it described the model.
Runs that were cut off mid-window are excluded from the board entirely β they are not counted as failures to solve. The completed run stands as a real result.
Three Moonshot Kimi models were entered and none produced a score. They were attempted four times and failed four different ways: one was killed by the host outage described above, one hit the $8 per-run spend ceiling at question 51, one was stopped after five consecutive replies that were not questions, and one was halted deliberately while we checked whether the provider was re-charging for text it already had.
Two things went wrong, and only one of them is about Kimi. Kimi repeatedly spends its entire output budget on reasoning and returns nothing at all β eight of sixteen turns in one measurement produced an empty reply, which still costs a full 6,000 tokens. Separately, the provider's prompt cache proved unreliable for this workload: hit rates swung between 95% and zero on consecutive turns of an identical, strictly-growing prompt, while DeepSeek served the very same prompts from cache almost perfectly. That is a provider-side routing effect, not something the benchmark can fix.
Kimi consumed $14.36 β more than half of everything we spent β and produced no completed run. Rather than spend more, it is recorded as a non-result with the transcripts kept in full. Giving Kimi a larger output budget would probably rescue it, but that would hand it more room to think than every other contestant got, so it would have to be a separate, clearly-labelled experiment rather than a 2.0 run.
Publishing this degrades it. Once "SaxAbench" and "Sax-a-Boom" sit next to each other on the open web, models will start recognising the setup and guessing sooner. That's why every run is date-stamped: a score is only meaningful relative to how poisoned the well was on the day it was recorded.