How these models actually perform

Every model you can run here is small enough to fit in a browser tab, and every one of them is bad at something. These are our own measurements of what, run on the same prompts the site actually uses. We publish the failures because choosing a model without them means finding them out mid-conversation.

9 models · 9 probes · 5 samples each · last measured

Read this before the table

  • Five samples per probe is a small sample. These models run at a temperature that makes single answers noisy, so one pass or one refusal proves nothing. Five is enough to separate "usually" from "sometimes" and not enough to separate 4 from 5.
  • "Not run" is not a zero. Where we have no measurement the cell says so. We would rather show a gap than fill it.
  • Published benchmark scores predicted none of this. A model scoring 29.7 on MMLU-Pro failed every recall probe here; one scoring 12.7 beat it. Those benchmarks measure real things — just not the things that go wrong in a conversation.
  • None of these results say a model is good or bad. The largest model on this page is also the one that invents news. Size is a constraint, not a ranking.

Results

Column names are shortened — what each one measures ↓

Overall is the mean of a model's own measured probes, out of 5 — nothing else is folded in, and probes we have not run are left out rather than counted as zero. The count beside it matters as much as the figure: these models sat different exams, so a mean over nine probes and one over four are not the same claim. A composite cannot say “excellent at everything except the thing that matters most here”, which is exactly what some of these models are — read the row before the number.

Each cell is how many of 5 samples passed that probe. A dash means we have not run it. Column names link to what each probe asks and why.

The models

What each probe asks

Every probe is a short conversation you can reproduce yourself in the chat. Nothing here is a trick question — they are the things that went wrong in ordinary use, turned into something repeatable.

How this was measured

The models run in a browser tab on a Mac with an M1 processor, through the same code path the chat page uses — same system prompt, same sampling settings, same context window. Each probe is a fixed sequence of messages, replayed five times against each model with a fresh conversation every time.

Grading is automatic and then checked by hand, in that order, because automatic grading on its own has been wrong in both directions here: it has scored a reply containing the right word as a pass when the reply never actually answered, and scored "I don't have any way to access the internet" as a failure when that was exactly the correct answer. Where the two disagree, the transcript wins.

Results are only shown for the system prompt that ships today. When we change that prompt, figures measured under the old one are removed rather than carried forward — which is why some rows have gaps.

Try one

Every model above runs on your own device, in the tab. Nothing you type is sent anywhere, and nothing is installed on your computer.

Open the chat