Are some types of AI image easier to recognise than others? The usual advice suggests they should be. Cars expose broken geometry. Food can look suspiciously perfect. Famous paintings give us an original composition to compare against.

We tested that assumption with real play data from WhichOneIsReal covering the last 28 days: 25 June–22 July 2026, inclusive. Players completed 356 comparable category-specific image games. Each game contained ten decisions, giving us 3,560 real-versus-AI choices across six categories.

The result is not a simple league table. The apparent differences may come from the subject category, the AI models used in that category, the particular images in the pool, or a combination of all three. The available analytics cannot separate those effects.

The results

The accuracy below is the mean of the accuracy value recorded when a game was completed. We excluded mixed mode because it combines categories, and excluded video, text and specialist modes because they test a different task. We also excluded the Faces mode because it presents two choices per round, while every category below presents four; its accuracy is therefore not directly comparable.

Example
Category
Accuracy
Completed games
AI alternatives in the current quiz pool
Famous paintings quiz illustration
Paintings
67.1%
55
30 AI images
Model names were not recorded for these images (10 legacy painting rounds).
Real sushi photograph from the food quiz
Food
60.2%
58
63 AI images
  • Stable Diffusion16
  • Nano Banana Pro13
  • GPT Image 1.59
  • FLUX family9
  • Kling5
  • DALL·E / MS Designer4
  • Seedream3
  • Z-Image1
  • Qwen Image Max1
  • Imagen 41
  • Lucid Realism1
FLUX family breakdown: 6 FLUX.2 Max, 1 FLUX Schnell, 1 FLUX Dev, 1 Juggernaut FLUX Pro.
Real group photograph from the people quiz
People
51.8%
67
30 AI images
  • FLUX family10
  • Kling5
  • Z-Image5
  • GPT Image 1.55
  • Seedream 4.55
FLUX family breakdown: 5 Juggernaut FLUX Pro, 5 FLUX.2 LoRA Gallery Realism.
Real tiger photograph from the animals quiz
Animals
49.4%
108
72 AI images
  • FLUX family17
  • Seedream 4.510
  • Z-Image7
  • GPT Image 1.55
  • Kling2
  • Nano Banana Pro1
42 of these are labelled by model — the other 30 have no saved label.
Real night-time car photograph from the cars quiz
Cars
36.5%
31
33 AI images
  • FLUX family15
  • GPT Image 1.54
  • Kling4
  • Z-Image3
  • Seedream 4.53
  • DALL·E 31
  • MS Designer1
  • Qwen Image Max1
  • Lucid Realism1
Real Taj Mahal photograph from the landmarks quiz
Buildings & places
35.9%
37
48 AI images
  • Imagen 229
  • DALL·E 35
  • Amazon Titan5
  • MS Designer3
  • Stable Diffusion3
  • FLUX Schnell1
  • FLUX Dev1
  • Lucid Realism1

How to read the model counts: "16 Stable Diffusion" means that 16 of the AI-generated images currently available in that category were made with Stable Diffusion. These are inventory counts, not correct answers, player impressions or the number of times Google Analytics recorded a model. Every round displays three AI alternatives and one real image. Because the export contains only category-level results, it does not reveal which of these images each player actually saw.

Why were buildings and cars so difficult?

Buildings and places produced the lowest accuracy: 35.9%. Cars followed at 36.5%. Both subjects share properties that favour current image generators.

Famous landmarks are photographed from a limited set of familiar viewpoints. Generators have seen countless images of the Eiffel Tower, Taj Mahal and Sydney Opera House, and they can reproduce the broad geometry, lighting and tourist-photo composition extremely well. A quiz player may know what the landmark looks like without knowing which particular photograph is authentic.

Cars have the same problem in a different form. Glossy paint, clean reflections, dramatic roads and showroom lighting are common in both advertising photography and AI training data. Synthetic cars can look more like an intentional photograph than a noisy, distant or poorly lit real image.

The model inventory also prevents us from assigning the result entirely to subject matter:

That concentration creates a confound. The low landmark result could reflect the category, Imagen’s ability to reproduce famous architecture, the particular rounds in the pool, or all three. The same applies to cars and their concentration of FLUX images. It is entirely possible that part — or even most — of the apparent category difference is actually a model or image-pool effect. We would need per-round answer data with model labels — and preferably the same models represented across every category — to separate those effects reliably.

Animals and people landed near 50%

Players scored 49.4% on animals and 51.8% on people. These categories offer more possible clues than landmarks, but many of the old shortcuts have weakened.

Modern generators reproduce fur, feathers, skin and individual faces convincingly. The remaining errors are often relational: a paw meeting the ground incorrectly, a hand gripping an object without believable pressure, inconsistent eye reflections, or several people interacting without a physically coherent pose. Those details require slower inspection than counting fingers — a tell that has largely stopped working.

The people pool is especially useful for understanding category effects because its model mix is relatively balanced. FLUX accounts for one third of its alternatives, while Kling, Z-Image, GPT Image and Seedream each account for one sixth. No single family dominates the category, yet accuracy remained close to 50%. That makes “complex human interaction is difficult to inspect” a more plausible explanation than it would be in a category controlled by one model.

The animal pool is less certain. Ten of its 24 rounds have no saved model labels, and FLUX represents 40.5% of the alternatives that are labelled. We can describe the observed difficulty, but not cleanly attribute it to either animals or models.

Why food and paintings were easier

Food reached 60.2% despite having the broadest recorded generator mix. Stable Diffusion was the largest contributor, but it represented only a quarter of the alternatives. Nano Banana, FLUX, GPT Image, Kling, DALL·E and several other systems were also present.

That diversity makes a pure model explanation less convincing. Food may offer category-specific clues: repeated garnishes, implausibly tidy crumbs, sauce with an overly smooth texture, duplicated ingredients and lighting that makes every part of the dish look equally perfect. Real food photography is styled too, but it still contains small accidents that generation tends to remove.

Paintings reached 67.1%, the best result among the four-choice categories. Familiarity probably helps. Many rounds use canonical works such as The Starry Night or Mona Lisa. Players may recognise the original composition, brushwork or proportions rather than detect a generic AI artifact.

There is also a hard limit to this interpretation: model labels were never stored for the ten legacy painting rounds. We cannot test whether their AI alternatives came from an easier or older generator set. Painting model attribution should be added before using this category in a model comparison.

What this study can — and cannot — tell us

This is behavioural data from real quiz sessions, not a controlled laboratory experiment.

The safest conclusion is therefore:

In this dataset, players struggled most with cars and famous places, while paintings produced the highest accuracy. However, the model mix differs substantially between categories. The data therefore shows an association, not proof that the category itself caused the difference.

What we will measure next

Aggregate game accuracy is enough to reveal a pattern, but not enough to explain it. A stronger follow-up should record anonymous round-level outcomes alongside:

With that data, we could compare the same category across models, the same model across categories, and individual rounds within each pool. We could also identify which specific visual situations fool players — not merely which menu category they came from.

For now, the most practical lesson is to slow down on architecture and vehicles. Do not rely on a landmark looking familiar or a car looking photographically polished. Check repeated structures, window and wheel geometry, reflections, contact with the road, background traffic and whether every perspective line belongs to the same scene — the relational checks from our guide to detecting AI images.

Put these numbers to the test — jump straight into the categories from this analysis:

Data and methodology