Scott Alexander’s AI Art Turing Test produced two results that are easy to confuse. First, roughly 11,000 participants achieved a median accuracy of 60% when deciding whether 50 images were human-made or AI-generated—only ten percentage points above chance. Second, when asked which image they liked best, participants often chose the machines: the two most popular works were AI-generated, as were six of the overall top ten.

The first result says AI art can be hard to identify in a deliberately curated challenge. The second is more awkward: selected images succeeded aesthetically even among respondents who reported a negative artistic opinion of AI art.

Those conclusions need caveats. This was a self-selected online audience, not a representative scientific sample, and Alexander deliberately curated away many easy giveaways. The experiment did not prove that AI art is always indistinguishable from human work—or that people prefer it in general. It did show how strongly judgments leaned on style stereotypes when origin labels and obvious defects were absent.

The AI Art Turing Test results at a glance

Alexander launched the challenge on Astral Codex Ten in October 2024 and published the full results in November. Participants saw 50 artworks and labelled each one human or AI.

A 60% median is better than guessing, but not by much. Across 50 decisions, the typical participant classified about 30 correctly instead of the 25 expected by chance.

Could you beat the 60% median? Try 10 rounds and find the authentic painting among three AI-generated alternatives. →

What did the test actually measure?

The phrase “Turing test” is an analogy to Alan Turing’s imitation game, not a formal certification that machines had “passed” art. Alexander adopted an often-cited reading of Turing’s 1950 prediction—treating 30% of judges being fooled as a pass—and noted that viewers misclassified the average AI image as human 40% of the time. By that loose analogy, AI art passed.

But this was not Turing’s original setup. There was no live exchange with an artist and no attempt to measure creativity as a process. The task was narrower: could a viewer infer an artwork’s origin from the finished image alone? There is no standardised pass threshold for that task.

Alexander initially planned five human and five AI works in each of four broad styles, for 40 images. Strong AI submissions led him to expand the set to 50. The human side included established artists such as Domenichino, Paul Gauguin and Jean-Michel Basquiat alongside digital artists; the synthetic side combined strong hobbyist submissions with other AI works Alexander curated from online galleries and creators.

That structure matters. This was closer to a best-case stress test than a random sample of images from the internet.

Why 60% is not an everyday detection rate

Alexander deliberately made the challenge difficult. He removed AI submissions with garbled text, malformed hands or other obvious defects. He also avoided human works with clues that would make their origin too easy to infer, such as complex interlocking poses, readable text and some forms of pop art. Familiar AI “house styles” were largely excluded too.

In other words, the test asked whether people could detect AI from subtle differences in style and quality after the cheap shortcuts had been removed. It did not ask whether people could catch the average low-effort AI image in a social feed.

There are other limits:

So “humans scored 60%” is meaningful. “Humans can only detect 60% of AI art” is not.

Style stereotypes fooled people

Participants were warned not to assume that oil paintings were human and digital images were AI. They still did it.

Answers clustered by visual style even though each style was close to evenly divided between human and synthetic works. Traditional-looking images attracted “human” labels. Polished digital art attracted “AI” labels. One striking example was Gauguin’s Entrance to the Village of Osny: the only genuinely human Impressionist work in the set was judged more artificial than the AI-generated Impressionist pictures.

Alexander’s normalised comparison made the bias concrete: respondents treated a near-even mix of 19th-century art as if 75% were human, while they treated a near-even mix of digital art as if only 31% were human. Mitchell Stuart’s genuinely human digital artwork Victorian Megaship was labelled AI-generated by 84% of respondents.

This reveals a common category error. People often detect the look associated with AI rather than the production process itself. When generators imitate an older medium, or a human artist works in a polished digital style, the shortcut reverses.

The same problem appears in everyday image detection. A cinematic glow is not proof of generation, and visible brushwork is not proof of a human hand. Better detection focuses on whether details relate coherently—lighting, reflections, contact points, repeated structure—not whether an image belongs to a style we mentally file under “AI.” Our guide to AI image tells that still work explains that shift in detail.

The uncomfortable result: AI art won the beauty vote

The accuracy score was not the experiment’s most provocative finding. Participants also selected a favourite image from the set. The two most popular images were AI-generated, and AI supplied six of the top ten.

That result needs the same care as the detection score. Many of the favourite AI images were Impressionist, and that category happened to include more AI work. When Alexander removed all Impressionist images from the ranking, human works reclaimed the top two places—but an AI image remained third, and AI still held four of the revised top ten.

The preference also crossed stated attitudes. Of the respondents, 33% reported a negative artistic opinion of AI art, 24% were neutral and 43% were positive. Among the 1,278 people who selected the most negative possible rating, the two most frequently chosen favourites were still AI-generated, and AI occupied five of their top ten.

That is not proof of hypocrisy. Someone can object to training practices, labour displacement, sameness or the flood of low-effort output while still liking a particular image. Alexander’s selection showcased strong, stylistically varied AI work and excluded much of the repetitive material critics encounter online.

The tension is still informative: among respondents who reported negative artistic opinions of AI art, AI images still ranked among the group’s most-selected favourites in the blind vote. That is an association within this survey, not evidence that removing a label caused anyone’s opinion to change.

A small group really could tell

The median result should not be mistaken for proof that nobody has a trained eye. Participants who said they disliked AI art averaged 64%, self-described professional artists averaged 66%, and respondents who were both professional artists and strongly anti-AI averaged 68%.

Five people scored 49 out of 50. Alexander cautioned that he could not separate luck from repeatable skill without retesting them, but results that high are extremely unlikely to come from random guessing alone.

These subgroup labels were self-reported, and Alexander did not publish confidence intervals or significance tests for the differences. They show an association, not proof that artistic training or disliking AI caused the higher scores.

The cautious lesson is not that AI origin is invisible. It is that detection skill is uneven and may depend on slow inspection of whether details have a coherent purpose. That is very different from relying on a generic “AI look.”

What the test changes—and what it does not

The AI Art Turing Test supports three modest conclusions:

  1. Origin is difficult to infer from style alone. A curated set of strong synthetic images pushed most participants close to chance.
  2. Aesthetic response and origin judgment are separate. Viewers could like an image while misclassifying who—or what—made it.
  3. The testing conditions control the headline. Remove obvious failures and select the best outputs, and AI looks stronger than it does in a random feed.

It does not settle whether AI-generated work is art, whether training was ethical, whether prompt authorship equals painting, or whether human artists remain socially valuable. A perception test cannot answer those questions.

The practical lesson is narrower: if authenticity matters, visual intuition should be treated as evidence, not proof. Check provenance, source and context. And if you want to improve your eye, practise on matched examples where the answer is revealed immediately—the feedback loop behind how we build our real-vs-AI quizzes.

Can you beat the 60% median?

Our AI art quiz uses famous paintings as the human originals and places them beside AI-generated imitations of the same subject. It is not a reproduction of Alexander’s survey: each round asks you to find one real masterpiece among three synthetic alternatives, and the content pool changes as models improve.

For a separate benchmark from our own players, see our AI image detection accuracy by category analysis of 3,560 quiz decisions. Paintings produced 67.1% observed accuracy, but its four-choice format, canonical artworks and different image pool mean that result is not directly comparable with Alexander’s binary test.

Six image categories beside an abstract green accuracy chart on a dark analytics dashboard.
Our category analysis found the highest observed accuracy for paintings, but it is a different test format—not a replication of Alexander's survey.

That makes it useful for the same underlying question. Once every option shares the same subject and obvious shortcuts disappear, can you recognise coherent composition, brushwork and detail—or are you relying on what “looks like AI”?

FAQ

What was the AI Art Turing Test?

Scott Alexander’s 2024 Astral Codex Ten survey asked roughly 11,000 people to classify 50 artworks as human-made or AI-generated. The images covered traditional, modern and digital styles.

What score did people get on the AI Art Turing Test?

The median participant scored 60%, while the mean was 60.6%. Because every image had two possible labels, blind chance would score 50%.

Did people prefer AI art to human art?

Within this curated set, the two most frequently selected favourite images were AI-generated and six of the top ten were AI-generated. That does not prove people prefer all AI art in general.

Did AI art pass the Turing test?

Alexander applied a commonly cited benchmark in which fooling 30% of judges counts as a pass, and AI exceeded it in his survey. But the experiment was not Turing’s original imitation game, and there is no standardised pass threshold for AI art.

Why was the AI Art Turing Test so difficult?

The organiser removed obvious AI failures such as malformed hands and garbled text, avoided human works with easy origin clues and selected strong examples from several styles. The test therefore focused on subtle differences rather than typical online images.

Were professional artists better at identifying AI art?

Self-described professional artists averaged 66%, compared with a 60.6% overall mean. The groups were self-reported within a self-selected sample, so the result shows an association rather than proof that artistic training caused the difference.

Sources and update policy

These results describe a fixed 2024 survey and do not change with current model releases. We checked the article against Alexander’s original challenge and results post on 4 August 2026. If the source publishes corrections or additional methodology, we will update this summary.