AI Art Won Blind Tests—Yet an AI Label Lowered Its Ratings

Selected AI-generated images can outperform human-made works when viewers judge pictures without knowing their origins. Subsequent controlled experiments have strengthened that narrow finding, while also showing that it does not amount to a general verdict that audiences value machine-made art more highly.
The apparent contradiction remains important for creators: people can prefer an image in a blind comparison and still downgrade it after learning that AI produced it. Aesthetic attraction, confidence about authorship and judgments about creative value measure different responses.
The blind test was striking, but deliberately selective
Scott Alexander’s November 2024 results covered an informal online challenge in which 11,000 participants classified 50 images as human- or AI-made. The median participant identified 60% correctly; the two most popular images were generated with AI, as were six of the top ten. Among the 1,278 respondents who gave AI art the lowest score on a five-point attitude question, AI images still occupied the first two favorite positions and five places in the top ten.
That pattern supports a specific claim: stated hostility to AI art did not reliably predict which pictures those respondents liked when authorship was concealed. It does not show that the respondents had abandoned concerns about training data, labor, originality or the place of artists in creative industries. Those issues were outside the visual-choice task.
The selection rules also limit how widely the result can be applied. Generated images with familiar defects, garbled writing or conspicuous model-specific aesthetics were excluded, while the human set avoided clues that would make authorship obvious. The AI entries were chosen from submissions by experienced hobbyists, so the test compared curated finalists rather than typical output from an image generator.
Style further complicated the ranking. AI-generated Impressionist pictures performed especially well, and that style was not evenly divided between the two categories. Once the Impressionist works were removed, human images moved into the first two positions, although generated work still appeared among the revised favorites. The outcome therefore reflects both origin and the particular mix of subjects and styles.
A controlled comparison supported the narrower result
More rigorous evidence arrived in a paper published on January 8, 2025. The Western Sydney University research paper paired 50 lesser-known representational artworks with 50 DALL·E 2 images designed to resemble their styles and visual characteristics. A group of 127 participants selected the image it preferred from each pair without being told that AI was the subject of the experiment, while a separate group of 137 attempted to identify the computer-generated image.
Participants in the preference experiment chose the generated works significantly more often than chance. The detection group also distinguished the AI images at an above-chance rate, and images that were easier to identify as synthetic tended to receive stronger preference scores. Their appeal therefore did not depend on consistently passing as human.
The controlled design makes this stronger evidence than an open online poll, but it retains clear boundaries. The participants were mainly young university students receiving course credit, the pictures were standardized and displayed digitally, and the generated set came from one model. Human selection was still involved because outputs that failed to match the predetermined comparison criteria were omitted.
Those conditions matter because digitization can suppress physical qualities such as scale, surface and texture. The experiment tested reactions to small representational images on screens, not encounters with original paintings or judgments across illustration, conceptual art, photography and other visual categories. It demonstrates a preference under defined conditions, not a universal hierarchy of artistic quality.
An authorship label changes the question
A second experiment isolated the effect of attribution instead of asking viewers to choose between human and generated images. In a paper published on May 12, 2025, the University of Notre Dame research team showed 92 undergraduates the same 24 DALL·E 2 images of biblical scenes. Half were truthfully told that AI had generated the works, while the others were told that art students had made them.
The human-attribution group evaluated the identical pictures more positively, including on likability, sincerity, composition, aesthetic value and emotional effect. Eye-tracking measures did not show reliable differences between the groups. The information about origin changed what participants said about the work without producing a comparable difference in the measured patterns of visual exploration.
This experiment does not prove that every AI disclosure will reduce audience response. It used religious subject matter, a modest student sample and a deliberately false attribution for one group. It does, however, demonstrate that the creator label can alter evaluation even when nothing in the image changes.
What the evidence means for creators
The defensible conclusion is that visual appeal and perceived artistic value are separable. A generated image may attract attention or win an unlabeled comparison, while its origin later affects whether viewers regard it as sincere, original or worthy of support. A click, favorite or dwell-time metric cannot capture all of those judgments.
The results also belong to complete production pipelines, not to image models acting independently. Someone chose the prompts, generated alternatives, rejected weak outputs and selected the images shown to participants. In the informal challenge especially, curation was part of the tested result.
For the creator economy, the central tension is therefore broader than whether AI can produce an attractive picture. Blind tests indicate that it can, under carefully selected conditions and even for some viewers who dislike AI art. Attribution experiments show that audiences may apply additional standards once authorship and process become visible—a distinction that matters wherever attention, reputation and payment depend on more than immediate visual preference.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.