How We Measure
Two kinds of number appear on this site: how often players identify the real item in a given category, and how convincing each AI model is. Both come out of answers given here, and both can be wrong in specific ways. This page says how they are computed and where they are weaker than they look, so you can decide how much weight to put on them.
What gets recorded
When you answer a round we store the answer, not you. One row holds the media type, the category, the full set of options you were shown, which one you picked, whether it was right, and how long you took. There is no account requirement and no name attached — the identifier is a random string generated in your browser, which exists only so that a repeated round can be recognised as a repeat.
Replays are excluded, not down-weighted
Category selection is seeded by date and mode, so everyone playing a given category on a given day gets the same ten rounds — and replaying that day means seeing the identical rounds again. A second look is not a fresh judgment about whether you can tell AI from a photograph; it is memory. So only your first answer to a round counts. Later ones are dropped entirely rather than counted at reduced weight, because a memory-assisted answer is not partial evidence — it is contaminated.
The content pool version is part of that key. If a round is edited — a new AI image swapped in for one of the slots — a previous look at the old picture says nothing about familiarity with the new one, so the round counts as fresh again.
Category accuracy, and why the baseline is not always 25%
Category accuracy is the share of rounds in which the real photograph or clip was the one picked. On its own that number is not comparable across modes, because they do not all offer the same number of options. Video rounds show exactly two clips, and so does the "which face is real" pair, so guessing there scores 50%. Ordinary image categories currently show four options per round, so guessing scores 25%.
This matters more than it sounds. Faces at about 70% look far easier than nature at about 41%, but faces are a two-image pair — so measured against their own baselines faces sit near 1.4× chance and nature near 1.6×, and the ordering reverses. Reading the raw percentages side by side tells you the opposite of what the data says. Every category figure on this site is therefore shown next to its own baseline and the ratio between them, and the baseline is read from the actual number of options in each set rather than assumed.
We publish a category only once it clears a floor of 300 rounds across at least 8 distinct real items. Below that the figure describes a handful of particular pictures rather than the category, so the page simply shows nothing.
Model ratings: a Bradley-Terry fit, anchored to reality
The model ratings answer a different question: how often is a given generator's output picked as "the real one"? Each round is treated as a top-one choice out of the field actually shown, and the model is fitted with the MM algorithm. Because the denominator is the field that appeared on screen, two-option and four-option rounds can be fitted together — their differing baselines are already accounted for. The same mechanism handles vote splitting: beating two strong rivals moves a model further than beating two weak ones.
Real photographs are the anchor, fixed at 1500. A model above 1500 is picked as the real one more often than genuine media is. Ratings are converted so that one point means what it means in chess, and the whole fit is recomputed nightly.
A model is only ranked once it has at least 10 distinct images and 30 appearances behind it. Thin models still appear on the ratings page by name, listed as not yet rated — removing them entirely would leave a reader who just met a generator in the quiz unable to find it, and conclude the table was broken.
Where these numbers are weakest
The confidence intervals are too narrow, and most so where it matters. Every round counts as one independent observation, but a model with ten images generates thousands of rounds out of those same ten pictures. The observations are clustered by image, so the effective sample size is far closer to the image count than to the round count — and a model measured on ten images gets an interval that looks more precise than one measured on a hundred. Do not read two overlapping ratings as a ranking, and do not read week-to-week movement as a trend; at our volumes that is noise with a headline. Fixing this properly needs a bootstrap that resamples images rather than rounds, which is scoped separately rather than patched with an arbitrary inflation factor.
Category and model effects are entangled. Not every model appears in every category in equal proportion, so a category that looks hard may be hard because of its subject matter or because the stronger generators happen to be over-represented in it. We cannot currently separate the two, and we do not claim to.
Our players are not a random sample. People who seek out an AI-detection quiz are more interested and probably more practised than the general public. Read these figures as describing our players, not humanity.
How the quizzes themselves are built
Sourcing, pairing and review — where the real media comes from, how fakes are matched to it, and when a pairing gets retired for being too easy — are covered separately in how we build the quizzes.
Corrections
If a number here looks wrong, it may well be — tell us and we will check it and say so publicly if it was. Write to info@whichoneisreal.com or use the contact page. Editorial responsibility for this site rests with its publisher, named in the legal notice.