The standard advice for spotting an AI image is to slow down. Look at the hands. Check the reflections. Zoom into the background. It sounds obviously right, and every guide on the subject repeats it β including ours.
So we checked it against our own play data. The result runs the other way.
The results
Every answer on WhichOneIsReal records how long the player took. We took 10,231 answers from four-image rounds β one real photograph, three AI-generated β and grouped them into two-second bands.
| Time taken | Answers | Picked the real photo |
|---|---|---|
| under 2s | 1,466 | 69.6% |
| 2β4s | 2,469 | 66.5% |
| 4β6s | 1,820 | 58.2% |
| 6β8s | 1,222 | 54.8% |
| 8β10s | 794 | 48.1% |
| 10β12s | 506 | 51.2% |
| 12β14s | 375 | 47.5% |
| 14β16s | 288 | 55.2% |
| 16β18s | 197 | 54.8% |
| 18β20s | 140 | 49.3% |
| over 20s | 954 | 54.2% |
How to read this: four images per round means pure guessing scores 25%. Every band is above chance β the question is by how much. The first five bands carry between 794 and 2,469 answers each; the bands past ten seconds carry 140 to 506, which is why they bounce around.
Accuracy falls from 69.6% to 48.1% across the first five bands, without a single reversal. Then it flattens. Everything past ten seconds sits around 52% and jumps up and down in a way that is fully explained by the smaller sample sizes.
So there are really two findings. Between zero and eight seconds, longer means worse β steadily. After eight seconds, more time changes nothing at all.
What this does not mean
It would be easy to turn this into βtrust your gut, answer fast.β That conclusion does not follow, and the data cannot support it.
The problem is direction. Two explanations fit these numbers equally well:
Time measures the image, not the thinking. An obvious fake is obvious immediately β you click fast and you are right. A well-made fake gives you nothing to grab onto, so you hesitate, study it, and pick wrong anyway. On this reading, response time is a symptom of difficulty, and the accuracy gap was decided before anyone started deliberating.
Deliberation actively hurts. You spot something in the first second, then talk yourself out of it.
We think the first is far more likely, and nothing here distinguishes them. A player who forces themselves to answer in one second on a hard image does not thereby become accurate β they just move themselves into a faster band while getting it wrong.
What it does mean
The useful part is the plateau, not the slope.
Past roughly eight seconds, the extra time buys nothing. If you have looked at an image for ten seconds and still cannot decide, the data says a further thirty seconds will not rescue you. That is a genuinely actionable finding, and it is the opposite of βkeep looking.β
It also fits what we found when we analysed accuracy by category: the reliable tells are relational β how light falls across a whole scene, whether an object actually touches the ground β and either you notice them or you do not. They are not the sort of detail that emerges from staring.
And it fits the shift described in the AI tells that stopped working. The old checklist rewarded slow inspection: count the fingers, read the sign, look for the melted background. Those checks are largely dead. What replaced them is closer to pattern recognition than to inspection β which is exactly the kind of judgement that arrives fast or not at all.
Method
The figures come from our own answer analytics, one row per decision, recorded anonymously with no name or email attached.
- Scope. Only four-image image rounds. Pair modes such as Which Face Is Real give a 50% baseline instead of 25%; mixing them in would produce an average belonging to neither. Video rounds are excluded for the same reason.
- Exclusions. Answers under 0.5 seconds (misfires, double taps) and over 60 seconds (abandoned tabs, backgrounded phones) are dropped. They are a small fraction of the total and would otherwise contaminate both ends.
- Bands. Two-second buckets up to 20 seconds, then a single band for everything slower. Bands below 100 answers are not reported.
- What is counted. Every recorded answer, including repeat plays by the same person. This is not a controlled experiment: our players are self-selected, more practised than the general public, and free to replay.
One limitation worth naming: we cannot see whether a player spent their ten seconds studying the image or looking at their phone. βTime to answerβ is not the same as βtime spent looking,β and the gap between them is invisible to us.
Put it to the test β the mixed round draws from every category, so it is the fairest way to time yourself: