Which AI model fools people most?
Live ranking from real quiz answers. Genuine photographs are fixed at 1500 — anything above that line gets picked as “the real one” more often than reality does.
Limited to the models built into this quiz — not every model on the market, and not always the latest version of each. A model's absence here says nothing about how good it is.
Image models
31 models listed · 846 AI images in the quiz · click a column to sort| Confidence | ||||||||
|---|---|---|---|---|---|---|---|---|
| — | Real image | 1500 | — | 46994 | 28676 | 61% | — | |
| 1 | Ideogram 4.0 QualityIdeogram | 1387 | 1365–1408 | 1296 | 335 | 26% | 10 | |
| 2 | Nano Banana 2 with Web SearchGoogle | 1381 | 1354–1408 | 859 | 219 | 25% | 10 | |
| 3 | Recraft V3Recraft | 1370 | 1345–1395 | 944 | 252 | 27% | 10 | |
| 4 | Imagen 2Google | 1357 | 1344–1370 | 6934 | 1340 | 46% | 29 | |
| 5 | Grok Imagine Image QualityxAI | 1353 | 1327–1380 | 951 | 213 | 22% | 10 | |
| 6 | Nano Banana 2 LiteGoogle | 1343 | 1318–1367 | 1162 | 247 | 21% | 10 | |
| 7 | Qwen Image 2.0 ProAlibaba | 1325 | 1301–1350 | 1172 | 236 | 20% | 10 | |
| 8 | Reve 2.1Reve AI | 1324 | 1300–1347 | 1325 | 271 | 20% | 10 | |
| 9 | Z-ImageAlibaba | 1287 | 1280–1294 | 17307 | 3010 | 18% | 132 | |
| 10 | Kling Image V3Kuaishou | 1284 | 1262–1307 | 1650 | 281 | 17% | 10 | |
| 11 | Juggernaut Pro FluxRunDiffusion | 1284 | 1272–1296 | 6289 | 1053 | 20% | 32 | |
| 12 | FLUX.2 [klein] 9BBlack Forest Labs | 1257 | 1228–1286 | 1238 | 184 | 19% | 14 | |
| 13 | Stable Diffusion 3.5 LargeStability AI | 1255 | 1236–1274 | 2433 | 375 | 15% | 10 | |
| 14 | Seedream 4.5ByteDance | 1248 | 1237–1258 | 8418 | 1233 | 15% | 33 | |
| 15 | Imagen 3 FastGoogle | 1243 | 1181–1305 | 237 | 35 | 15% | 10 | |
| 16 | FLUX.2 [pro]Black Forest Labs | 1235 | 1202–1268 | 957 | 123 | 13% | 10 | |
| 17 | Stable DiffusionStability AI | 1232 | 1217–1248 | 4643 | 601 | 18% | 37 | |
| 18 | Gemini 2.5 Flash Image / Nano BananaGoogle | 1231 | 1214–1247 | 3784 | 510 | 13% | 14 | |
| 19 | FLUX.2 [max]Black Forest Labs | 1212 | 1188–1236 | 1936 | 231 | 12% | 10 | |
| 20 | Imagen 4 UltraGoogle | 1201 | 1164–1239 | 860 | 93 | 11% | 10 | |
| 21 | FLUX.1 [schnell]Black Forest Labs | 1183 | 1175–1191 | 21726 | 2130 | 10% | 157 | |
| 22 | Kling IMAGE 3.0 OmniKuaishou | 1182 | 1164–1200 | 4033 | 410 | 10% | 24 | |
| 23 | GPT Image 1.5OpenAI | 1176 | 1161–1191 | 5729 | 564 | 10% | 23 | |
| 24 | SDXL-LightningByteDance | 1172 | 1158–1185 | 7969 | 720 | 9% | 72 | |
| 25 | FLUX.1 [dev]Black Forest Labs | 1166 | 1153–1180 | 8237 | 748 | 9% | 41 | |
| 26 | MAI-Image-2.5Microsoft AI | 1166 | 1122–1211 | 730 | 64 | 9% | 10 | |
| 27 | Juggernaut Flux LightningRunDiffusion | 1151 | 1122–1179 | 1794 | 159 | 9% | 10 | |
| 28 | DALL·E 3OpenAI | 1148 | 1122–1173 | 2634 | 195 | 7% | 27 | |
| 29 | Seedream 4ByteDance | 1138 | 1114–1162 | 2634 | 215 | 8% | 13 | |
| 30 | Seedream 5.0 ProByteDance | 1090 | 1006–1174 | 276 | 17 | 6% | 10 | |
| 31 | Qwen-Image-MaxAlibaba | 1084 | 1051–1117 | 1842 | 113 | 6% | 12 |
Dimmed rows have fewer than 60 appearances — early evidence, not yet a settled rating.
A model joins this table once it has at least 10 images in the quiz and has appeared in at least 30 answers — below that, the number would say more about one lucky (or unlucky) image than about the model itself. New models and images are added regularly, so more entries — and firmer ratings — are on the way.
In the quiz but not yet listed: FLUX.2 LoRA Gallery: Realism (fal), FLUX1.1 [pro] (Black Forest Labs), GPT Image 2 Medium (OpenAI), Imagen 4 (Google), Lucid Realism (Leonardo AI), Titan Image Generator G1 v1 (Amazon) — fewer than 10 images in the quiz so far.
Last recomputed 18 Sept 2026, 03:15 · refreshed nightly
Add your own answers
Every round you play feeds straight into these numbers.
🎬 Video quizzes
Not in the ranking above yet — too few answers so far to rate video models.
How the rating works
Bradley-Terry, not a win count. Each round counts as a choice among the options actually on screen, so beating strong rivals is worth more than beating weak ones. Two-option rounds and four-option rounds therefore combine without their different odds distorting anything.
Read the range, not the position. Where two intervals overlap, the order between those models is not established. Dimmed rows have fewer than 20 appearances and are listed for completeness only.
Pool size is context. A model represented by three images is judged on those three images, however often they were shown.
Replays don't get extra weight — within a visit. Only your first look at a round counts, because after that you'd recognise it and it stops being a fresh judgment. That rule ends when you close the tab: come back later the same day and a repeat is counted again.
Limits. The method assumes all players judge alike, and this measures our image pool rather than a model's ceiling. Answers recorded before 29 August 2026 come only from players who accepted analytics cookies; since then every player counts, so figures spanning that date mix two different groups.


