Which AI model fools people most?
Live ranking from real quiz answers. Genuine photographs are fixed at 1500 — anything above that line gets picked as “the real one” more often than reality does.
Limited to the models built into this quiz — not every model on the market, and not always the latest version of each. A model's absence here says nothing about how good it is.
Image models
29 models listed · 849 AI images in the quiz · click a column to sort| Confidence | ||||||||
|---|---|---|---|---|---|---|---|---|
| — | Real image | 1500 | — | 4485 | 2455 | 55% | — | |
| 1 | Nano Banana 2 with Web SearchGoogle | 1500 | 1422–1578 | 80 | 31 | 39% | 10 | |
| 2 | Ideogram 4.0 QualityIdeogram | 1463 | 1398–1529 | 125 | 40 | 32% | 10 | |
| 3 | GPT Image 2 MediumOpenAI | 1417 | 1349–1484 | 125 | 35 | 28% | 10 | |
| 4 | Nano Banana 2 LiteGoogle | 1382 | 1302–1462 | 105 | 23 | 22% | 10 | |
| 5 | Grok Imagine Image QualityxAI | 1371 | 1260–1481 | 57 | 12 | 21% | 10 | |
| 6 | Reve 2.1Reve AI | 1345 | 1221–1468 | 52 | 9 | 17% | 10 | |
| 7 | Qwen Image 2.0 ProAlibaba | 1337 | 1250–1423 | 96 | 19 | 20% | 10 | |
| 8 | Juggernaut Pro FluxRunDiffusion | 1334 | 1298–1371 | 583 | 116 | 22% | 33 | |
| 9 | Z-ImageAlibaba | 1330 | 1308–1352 | 1579 | 307 | 20% | 133 | |
| 10 | Gemini 2.5 Flash Image / Nano BananaGoogle | 1325 | 1280–1371 | 347 | 70 | 20% | 14 | |
| 11 | Imagen 2Google | 1315 | 1279–1351 | 937 | 160 | 41% | 29 | |
| 12 | FLUX.2 [klein] 9BBlack Forest Labs | 1308 | 1208–1408 | 93 | 16 | 23% | 14 | |
| 13 | Stable Diffusion 3.5 LargeStability AI | 1296 | 1232–1360 | 199 | 34 | 17% | 10 | |
| 14 | Seedream 4.5ByteDance | 1295 | 1264–1326 | 871 | 145 | 17% | 33 | |
| 15 | FLUX.2 [pro]Black Forest Labs | 1282 | 1172–1392 | 75 | 11 | 15% | 10 | |
| 16 | Kling Image V3Kuaishou | 1281 | 1201–1362 | 138 | 21 | 15% | 10 | |
| 17 | FLUX.1 [schnell]Black Forest Labs | 1272 | 1250–1295 | 1920 | 273 | 14% | 157 | |
| 18 | Stable DiffusionStability AI | 1261 | 1210–1311 | 422 | 57 | 18% | 37 | |
| 19 | GPT Image 1.5OpenAI | 1257 | 1217–1296 | 651 | 87 | 13% | 23 | |
| 20 | Juggernaut Flux LightningRunDiffusion | 1253 | 1177–1329 | 173 | 23 | 13% | 10 | |
| 21 | SDXL-LightningByteDance | 1241 | 1197–1285 | 582 | 68 | 12% | 72 | |
| 22 | Seedream 4ByteDance | 1241 | 1162–1320 | 165 | 21 | 13% | 13 | |
| 23 | FLUX.1 [dev]Black Forest Labs | 1221 | 1180–1261 | 724 | 79 | 11% | 41 | |
| 24 | DALL·E 3OpenAI | 1203 | 1137–1269 | 299 | 29 | 10% | 27 | |
| 25 | FLUX.2 [max]Black Forest Labs | 1195 | 1125–1266 | 256 | 26 | 10% | 10 | |
| 26 | Kling IMAGE 3.0 OmniKuaishou | 1191 | 1137–1244 | 458 | 45 | 10% | 24 | |
| 27 | Imagen 4 UltraGoogle | 1186 | 1075–1297 | 113 | 10 | 9% | 10 | |
| 28 | MAI-Image-2.5Microsoft AI | 1156 | 988–1325 | 61 | 4 | 7% | 10 | |
| 29 | Qwen-Image-MaxAlibaba | 1101 | 997–1206 | 180 | 11 | 6% | 12 |
Dimmed rows have fewer than 60 appearances — early evidence, not yet a settled rating.
A model joins this table once it has at least 10 images in the quiz and has appeared in at least 30 answers — below that, the number would say more about one lucky (or unlucky) image than about the model itself. New models and images are added regularly, so more entries — and firmer ratings — are on the way.
In the quiz but not yet listed: FLUX.2 LoRA Gallery: Realism (fal), FLUX1.1 [pro] (Black Forest Labs), Imagen 4 (Google), Lucid Realism (Leonardo AI), Titan Image Generator G1 v1 (Amazon) — fewer than 10 images in the quiz so far; Imagen 3 Fast (Google), Recraft V3 (Recraft), Seedream 5.0 Pro (ByteDance) — shown too rarely so far, under 30 answers.
Last recomputed 10 Aug 2026, 03:15 · refreshed nightly
Add your own answers
Every round you play feeds straight into these numbers.
🎬 Video quizzes
Not in the ranking above yet — too few answers so far to rate video models.
How the rating works
Bradley-Terry, not a win count. Each round counts as a choice among the options actually on screen, so beating strong rivals is worth more than beating weak ones. Two-option rounds and four-option rounds therefore combine without their different odds distorting anything.
Read the range, not the position. Where two intervals overlap, the order between those models is not established. Dimmed rows have fewer than 20 appearances and are listed for completeness only.
Pool size is context. A model represented by three images is judged on those three images, however often they were shown.
Replays don't get extra weight. Only your first look at a round counts — after that you'd recognise it, and it stops being a fresh judgment.
Limits. Only answers from players who accepted analytics are counted, and the method assumes all players judge alike. This measures our image pool, not a model's ceiling.


