Which AI model fools people most?
Live ranking from real quiz answers. Genuine photographs are fixed at 1500 — anything above that line gets picked as “the real one” more often than reality does.
Limited to the models built into this quiz — not every model on the market, and not always the latest version of each. A model's absence here says nothing about how good it is.
Image models
19 models listed · 849 AI images in the quiz · click a column to sort| Confidence | ||||||||
|---|---|---|---|---|---|---|---|---|
| — | Real image | 1500 | — | 1898 | 1110 | 58% | — | |
| 1 | FLUX.2 [klein] 9BBlack Forest Labs | 1383 | 1258–1508 | 46 | 11 | 29% | 14 | |
| 2 | Kling Image V3Kuaishou | 1314 | 1196–1432 | 54 | 10 | 19% | 10 | |
| 3 | Z-ImageAlibaba | 1299 | 1263–1334 | 626 | 113 | 18% | 133 | |
| 4 | Imagen 2Google | 1295 | 1245–1346 | 494 | 78 | 38% | 29 | |
| 5 | Gemini 2.5 Flash Image / Nano BananaGoogle | 1295 | 1226–1363 | 165 | 30 | 18% | 14 | |
| 6 | Juggernaut Pro FluxRunDiffusion | 1291 | 1232–1349 | 253 | 43 | 19% | 33 | |
| 7 | Stable Diffusion 3.5 LargeStability AI | 1275 | 1169–1380 | 79 | 12 | 15% | 10 | |
| 8 | Seedream 4.5ByteDance | 1273 | 1227–1319 | 420 | 66 | 16% | 33 | |
| 9 | DALL·E 3OpenAI | 1263 | 1172–1354 | 121 | 16 | 13% | 27 | |
| 10 | Seedream 4ByteDance | 1263 | 1166–1361 | 99 | 14 | 14% | 13 | |
| 11 | SDXL-LightningByteDance | 1239 | 1149–1329 | 131 | 16 | 12% | 72 | |
| 12 | Qwen-Image-MaxAlibaba | 1239 | 1104–1374 | 54 | 7 | 13% | 12 | |
| 13 | Juggernaut Flux LightningRunDiffusion | 1233 | 1136–1329 | 112 | 14 | 13% | 10 | |
| 14 | FLUX.1 [schnell]Black Forest Labs | 1233 | 1192–1273 | 658 | 81 | 12% | 157 | |
| 15 | Stable DiffusionStability AI | 1228 | 1149–1307 | 191 | 22 | 15% | 37 | |
| 16 | GPT Image 1.5OpenAI | 1212 | 1149–1276 | 283 | 32 | 11% | 23 | |
| 17 | FLUX.1 [dev]Black Forest Labs | 1189 | 1123–1255 | 301 | 29 | 10% | 41 | |
| 18 | Kling IMAGE 3.0 OmniKuaishou | 1169 | 1085–1252 | 201 | 18 | 9% | 24 | |
| 19 | FLUX.2 [max]Black Forest Labs | 1103 | 964–1242 | 100 | 6 | 6% | 10 |
Dimmed rows have fewer than 60 appearances — early evidence, not yet a settled rating.
A model joins this table once it has at least 10 images in the quiz and has appeared in at least 30 answers — below that, the number would say more about one lucky (or unlucky) image than about the model itself. New models and images are added regularly, so more entries — and firmer ratings — are on the way.
In the quiz but not yet listed: FLUX.2 LoRA Gallery: Realism (fal), FLUX1.1 [pro] (Black Forest Labs), Imagen 4 (Google), Lucid Realism (Leonardo AI), Titan Image Generator G1 v1 (Amazon) — fewer than 10 images in the quiz so far; FLUX.2 [pro] (Black Forest Labs), GPT Image 2 Medium (OpenAI), Grok Imagine Image Quality (xAI), Ideogram 4.0 Quality (Ideogram), Imagen 3 Fast (Google), Imagen 4 Ultra (Google), MAI-Image-2.5 (Microsoft AI), Nano Banana 2 Lite (Google), Nano Banana 2 with Web Search (Google), Qwen Image 2.0 Pro (Alibaba), Recraft V3 (Recraft), Reve 2.1 (Reve AI), Seedream 5.0 Pro (ByteDance) — shown too rarely so far, under 30 answers.
Last recomputed 4 Aug 2026, 03:15 · refreshed nightly
Add your own answers
Every round you play feeds straight into these numbers.
🎬 Video quizzes
Not in the ranking above yet — too few answers so far to rate video models.
How the rating works
Bradley-Terry, not a win count. Each round counts as a choice among the options actually on screen, so beating strong rivals is worth more than beating weak ones. Two-option rounds and four-option rounds therefore combine without their different odds distorting anything.
Read the range, not the position. Where two intervals overlap, the order between those models is not established. Dimmed rows have fewer than 20 appearances and are listed for completeness only.
Pool size is context. A model represented by three images is judged on those three images, however often they were shown.
Replays don't get extra weight. Only your first look at a round counts — after that you'd recognise it, and it stops being a fresh judgment.
Limits. Only answers from players who accepted analytics are counted, and the method assumes all players judge alike. This measures our image pool, not a model's ceiling.


