Figure A1
Model: FLUX-pro (Black Forest labs)
Prompt: “Create a grainy TV quality photograph of Donald Trump talking at a political rally, supporters with trump 2024 signs behind him”
Declan Wright, May 17, 2025
History
ChatGPT’s release in November 2022 marked a rush of interest in AI development, capturing widespread public and academic attention. It sparked what many consider the next age of internet innovation, with massive economic opportunity for those involved. What started out as a research project evolved into a commercial product embraced by millions worldwide, surpassing initial projections of growth. It is now the seventh most used website globally, ranking only behind the likes of Google, Instagram and Facebook (Similarweb, 2025). Technologies powering ChatGPT and other tools are collectively known as generative artificial intelligence (AI). These models can create novel content that replicates human writing, speech, and visual media with realism. A foundational innovation leading to the generative models used today was conceptualized in the 1980s following simultaneous independent discoveries of the multilayer perceptron, or MLP (Rumelhart et al., 1986). This discovery is what allows AI models to learn patterns. They work by updating internal parameters based on observed examples. MLPs are loosely modeled after the human nervous system, but are primarily mathematical and statistical models designed for nonlinear pattern recognition. It is notable that while humans typically require one or two examples to learn something (Holzinger et al., 2023), training powerful neural networks requires massive amounts of data. Stable Diffusion, a popular open source image generator, was trained on 2.3 billion images (Baio, 2022).
In 2014, the first landmark AI technology was released for image generation. Known as Generative Adversarial Networks, or GANs, these models were a novel framework proposed by Goodfellow and colleagues at the University of Montreal (2014). GANs are trained through an adversarial process between two neural networks: a generator that creates images, and a discriminator that tries to distinguish between real and generated images. This setup creates a competitive dynamic where the generator improves at creating more realistic images to fool the discriminator, while the discriminator becomes better at detecting fakes. As an early innovation, GANs often struggled with training stability and model collapse. They also have other limitations. StyleGAN, a model developed by Nvidia in 2018, can only generate images of faces. Users cannot specify attributes like age, ethnicity, or expression, which restricts its flexibility and makes it less adaptable compared to newer models which can generate a wide variety of images (Karras et al., 2019).
Diffusion Models
Diffusion models, introduced more recently, have largely superseded GANs as the primary technology for image generation. The mathematical foundation for diffusion was conceptualized by researchers at UC Berkeley (Sohl-Dickstein et al., 2015). Adapted to image training, diffusion models are trained by adding fixed amounts of random noise to training images. In photography, noise refers to random pixels that disrupt the structure of an image. Diffusion models learn to reverse this noise addition process, enabling them to start with completely random pixels and progressively remove noise by predicting and subtracting what the noise should look like at each step of the generation process. This approach offers several advantages over GANs. Diffusion models are more stable during training and don't suffer from the same model collapse issues. They also can generate a wider variety of image concepts with better consistency (Dhariwal & Nichol, 2021). Additionally, the introduction of the CLIP encoder by OpenAI enabled diffusion models with even greater flexibility, allowing text prompting, more precise control over the generation process, and features like image editing (Ramesh et al., 2022). This enhanced controllability has led to explosive adoption. An estimated 34 million AI images were generated daily as of 2023, with approximately 15 billion AI images created by that time (Valyaeva, 2023). This number has grown exponentially since then. Following the release of a new image generation model in ChatGPT, the platform saw over a million new users sign up for the product in a single hour (Altman, 2025). While Instagram hosts approximately 50 billion images accumulated over more than a decade, AI-generated imagery has now likely exceeded that figure. Traditional stock photography platforms have been completely outpaced. Shutterstock's entire library of 386 million images represents less than a month's worth of current AI image production (Valyaeva, 2023).
Outlook and Concerns
The newfound popularity of image models has created significant risks. The findings presented in this study aim to evaluate the level of concern that may be appropriate for the widespread distribution and use of advanced image models. One of the most widely cited concerns with AI models is the potential for misinformation. During the 2024 US presidential election, prominent figures including Candidates in the 2024 US presidential election repeatedly posted AI-generated photos of other candidates and celebrities such as Taylor Swift, a newfound concern with voter influence (Robins-Early, 2024). Extensive wildfires in Los Angeles were almost immediately followed by a flood of AI-generated images online, depicting iconic landmarks like the Hollywood sign on fire, when they were not at risk from the fires (Chappell, 2025). More recently, similar engagement baiting content generated by AI depicting false scenes of earthquakes in Myanmar received millions of cumulative views online (de Graaf & Muros, 2025). In the past, creating manipulated images required extensive technical skill. Today, AI models allow images to be created with little skill, taking seconds to process. A 2024 preprint conducted by Stanford researchers found a significant wave of engagement-baiting content generated by AI on Facebook. This content receives hundreds of millions or engagements, with the majority of users not realizing they were viewing artificial content. Perhaps more concerningly, the Facebook algorithm appears to be actively promoting these posts, increasingly presenting them to a wider range of audiences (DiResta & Goldstein, 2024). This raises significant concerns about the ethical implications of algorithmic amplification, particularly how it enables the spread of misinformation and its potential to erode public trust in media. In publicly released community guidelines, Facebook itself, as well as other social media platforms including TikTok and X, all have similar policies on AI content: placing the burden on users to self-report AI usage (Meta, 2025; TikTok, 2024; X, 2025). Consequently, regulating AI misinformation will likely require a comprehensive approach. A meta-review of AI visual content concluded that effective regulation demands both watermarked credentials on photos and the deployment of AI detection systems working in tandem, as neither method alone represents a foolproof approach (Davis, 2024). Another promising approach when it comes to combating misinformation is education. Even basic knowledge of misinformation can help people improve their ability to identify it, despite gains remaining under 10% improvement (Ali & Qazi, 2021). This idea is consistent with historical trends as well. Even in ancient societies, human populations adapt to new forms of deception extremely quickly (Levine, 2020).
Image generators face other concerns. The usage of billions of images in training diffusion models has led to widespread outrage among artists and photographers whose work may have been used in model training. Highly publicized lawsuits, such as the 2023 challenge reported by Brittain (2023), have brought further awareness to this concern. While the issue of copyright is yet to be decided in the courts, some companies have taken a proactive approach to training their models. Adobe, the company behind popular image editing software such as Photoshop, stated in a 2024 release that its Firefly models were trained by paying artists for rights to use their images, and allowing users to have their images opted out of AI training (Adobe, 2025). Another issue has been the introduction of racial biases. Models trained on images that lean heavily towards a specific demographic group, or are tuned to produce output based on race, can have notable consequences. Following the release of Google’s flagship AI product Gemini, the company was forced to release a public apology and stop image generation features on the platform, after the model generated racially inappropriate images of certain historical groups (Raghavan, 2024). This incident not only sparked widespread criticism from the public but also raised concerns within the tech industry about the lack of rigorous testing for bias in AI models. It pressured Google to re-evaluate its training datasets and slow the rollout of similar features in subsequent projects. This feature was restored with the release of Imagen 3, Google’s latest model, which featured extensive de-biasing within training data and in tuning before being released (Baldridge et al., 2024).
While the risks associated with AI image generation are notable, the technology also offers tangible benefits. In the case of image generators, previous academic analysis of their impacts has mostly focused on economic benefits, with generative AI predicted to add trillions of dollars per year to the global economy (Chui et al., 2023). Investment in AI technologies has seen incredible interest, even during a period of slow general economic growth (Chui et al., 2023).
For creative professionals and designers, the use of AI is democratizing. For those without the previous knowledge and ability to create visuals, AI tools can bring ideas into the world. This concept also brings promise to marketing and prototyping. Nike worked with a group of elite athletes who were able to visualize prototypes and concepts using AI tools. This significantly accelerated the development process and allowed the Nike design team to bring athletes’ visualizations to reality (Creating the unreal, 2024).
Identifying AI-Generated Images
Previous research on an individual’s ability to differentiate between AI-generated images and real photographs is severely lacking. A Serbian study (Vukojičić et al., 2023) conducted by student researchers was the first directly presenting research on this topic in early 2023. The authors compared how well people could tell the difference between images created using DALL-E 2 and hand drawn sketches. This research found that performance varied greatly depending on the difficulty of the task, and whether images were presented in groups or in pairs. While it provides a good basis for understanding the topic, its sample was limited to a small group or students, and the method of presenting image pairs introduces bias that could have influenced the ability of participants to discern AI-generated drawings. More recent research involving collaborators across several institutions from across China, Hong Kong, and Sydney, evaluated human performance at discerning AI-generated images, and compared it to a model trained on the Fake2M dataset (Lu et al., 2023). In this research, humans were able to identify AI images correctly 61.3% of the time. The images in that dataset were created using Stable Diffusion, StyleGAN3, and DeepFloyd IF, a multi-process model. This does present some issues about the generalizability of the findings, as models like GANs and some forms of auto-regression are considered antiquated technologies. This directly leads to the development of a research question: “How well can individuals distinguish between photographs created using leading-edge generative diffusion models as compared to naturally created photographs, and how does providing feedback on their initial performance impact their confidence and ability to improve in subsequent attempts?” This guiding question aims to address these gaps in previous research, by using leading-edge diffusion models such as Imagen 3 and the FLUX family of models. Compared to generators used in previous research, such as Stable Diffusion and Dall-E 2, these leading-edge models bring significant improvements in image quality and photorealism, representing an area that remains understudied. This evaluation also addresses two other significant gaps in research. Firstly, as previous research has found positive effects on people’s ability to identify misinformation, a feedback mechanism will be implemented, allowing participants to observe their mistakes and potentially learn from them. Secondly, participants’ confidence will be measured at three points during the research process. Unlike previous studies which took the ability of participants at face value, this provides an additional dimension for interpreting the results.
This study aims to determine whether individuals can accurately identify AI-generated images. To achieve this, a dataset comprising photorealistic images generated by AI alongside real photographs was compiled, and a human survey using the images was conducted. The following sections explain the specifics of the data collection process, the human evaluation, and the techniques employed to analyze the findings.
Image Dataset
A total of 60 images were compiled to create the dataset used in this research. The images consist of 30 real photographs and 30 synthetic images, generated using advanced diffusion models.
Table 1
Models Used in Image Generation
| Model | Developer | Release Date | Images Generated (%) |
|---|---|---|---|
| Imagen 3 | Google DeepMind | May 14, 2024 | 31% |
| FLUX.1 Pro | Black Forest Labs | August 1, 2024 | 35% |
| FLUX 1.1 Pro | Black Forest Labs | October 2, 2024 | 17% |
| Recraft V3 | Recraft AI | October 30, 2024 | 17% |
After a diverse range of images were generated, the highest quality images were selected for use as a part of this research. This selective process does not result in random samples of generated images, but was chosen to mirror real-world environments where AI-generated images are optimized for “believability” and potential misinformation. Examples of images that were not used include images with obvious artifacts (a hand with ten fingers), or a lack of fine textures or general realism. To put this in perspective, for every image used in the dataset, an average of 4.83 additional images were generated and not used.
To prevent participants from making direct comparisons that could bias their decisions, the real and synthetic image datasets do not contain image pairs (e.g., there is not a synthetic image of a black cat paired with a real photograph of a black cat). Authentic photographs were sourced from reputable, free-use websites including Wikimedia Commons, Pixabay, and Flickr. All real images were published before 2021, as this predates the availability of modern AI image generators. To avoid bias in selecting images, the photographs are distributed evenly across six different categories with five images each. These categories are designed to represent a diverse array of image depictions across different visual regions, to ensure that there is not a strong leading or bias in the image data set (Lu et al., 2023). The use of any graphic, suggestive, or other sensitive images was avoided in order to ensure participants’ safety and comfort with the evaluation. Full details regarding each image, including file names, sources, AI model details (including prompts), visual categories, and resolutions, are provided in Appendix A.
In order to evaluate perceptual qualities of images, a texture and complexity rating scale from one to ten was created. Images with unique complexities or strong textures were given higher scores closer to ten, whereas images with low levels of texture or specific details were given ratings closer to one. These ratings were assigned by manually coding the images.
Evaluation Design
Images are divided randomly into two groups, each in separate survey sections. Google Forms was used to build and host the evaluation. Participants were recruited through community outreach, with a total of n = 280 individuals completing the evaluation. After the first set of images was evaluated, feedback with the correct answers to these initial photos was given to participants. The second set of images was evaluated under the same method as the first. Basic demographic questions were also asked for the purpose of analyzing differences across such demographic groups, specifically looking at gender, age, education level, location of residence, screen time, and a self-reported question about any factors that might enhance a participant’s ability (for example, artists, or people with experience using AI). Participants’ confidence levels were measured using three quantitative questions ranging from 1 to 6. This scale was chosen as it pushes participants to a certain result instead of allowing them to reply with a neutral answer, which is a common issue in survey design. The first question is placed before the initial image set, the second after reviewing the feedback from the first set, and the final question after the completion of the second image set. This approach is intended to show changes in participants’ confidence, which potentially offers an additional way to interpret ability beyond accuracy. Informed consent was included as a part of the evaluation. The complete instructions provided to participants, details about the feedback content, and the exact wording of the demographic and Likert scale questions can be found in Appendix B.
Statistical Analysis
Following the collection of data, the results were analyzed using Google Sheets. Data analysis was conducted in several ways. Initially, descriptive statistics were employed. Raw data from the evaluation was converted into numerical values. Correct responses were graded with a 1, and incorrect responses received a 0. Each participant’s overall accuracy was found by averaging these results and data was described using standard deviations.
Subsequently, a “random guessing” model was created. To build the random model, 30 random integers (either 0 or 1) were generated, and the mean of these integers was taken to represent a simulated participant’s score. This process was repeated 280 times to create a normal distribution of simulated participants with the same sample size as the real evaluation.
Statistical analysis methods, such as sample t-tests, were used to determine if the difference between the two image sets and the random model were statistically significant. A t-test is a statistical procedure used to determine whether the mean difference between two observations is convincing evidence to reject a null hypothesis. In this case, t-tests were used to compare participant accuracy before and after feedback, and were also used to compare confidence between individual questions.
Additionally, Analysis of Variance (ANOVA) was used to compare means across multiple categories simultaneously. Unlike t-tests which can only compare two groups, ANOVA allows for a similar comparison of performance across more than two groups. This method helped identify if there were statistically significant differences in accuracy among image categories and at the individual image level.
Finally, a regression analysis was used to determine the relationships between scores and confidence, and scores and demographic information. Statistical significance was modeled using the coefficient of determination R², which measures the percentage of variance in the data explained by the model. In addition, a multilayer perceptron (MLP) neural network was trained to model participant accuracy based on demographic factors and confidence ratings. The MLP consisted of two hidden layers (16 and 8 neurons) with rectified linear unit (Leaky ReLU) activation functions. See Appendix D for the full training script with comments explaining its precise implementation. This approach allowed for discovery of non-linear relationships that might not be apparent through standard regression analysis.
This statistical approach effectively addressed the research question by directly measuring participants’ ability to distinguish between AI-generated and real images, the effectiveness of feedback in enhancing detection skills, and changes in confidence levels and their correlation with performance.
The following section presents research findings, focusing on participant accuracy, changes in confidence levels, and the relationships between these variables and demographic factors.
Overall Accuracy
Participants' ability to correctly identify AI-generated images was assessed across two sets of images. The mean accuracy for image set one, before feedback, was 48.7%. For image set two, which participants completed after receiving feedback, the mean was 48.3%. A paired-sample t-test resulted in a p value of 0.627, which is not convincing evidence for statistically significant difference between the accuracy rates of the two image sets.
Furthermore, participants' performance was compared to a random guessing model. Sample t-tests indicated that accuracy for both image sets was not statistically different from chance. This suggests that, on average, participants were unable to reliably distinguish between real and AI-generated images, even after receiving feedback.
Table 2
Descriptive Statistics for Accuracy
| Image Set | Mean Accuracy (%) | Standard Deviation (%) | p-value (vs. random model) |
|---|---|---|---|
| Set 1 (before feedback) | 48.7% | 10.1% | 0.398 |
| Set 2 (after feedback) | 48.3% | 9.6% | 0.876 |
Note: A standard significance level of α = 0.05 was used to reject the null hypothesis.
Individual Image Analysis
While overall accuracy was near chance levels, a one-way ANOVA revealed highly statistically significant differences in correct identification rates between individual images. This resulted in an F statistic of approximately 51.41 with 59 degrees of freedom between groups and 16,770 degrees of freedom within groups. Accuracy rates ranged from a low 12.1% (for an AI-generated image of a woman painter) to a high of 88.6% (for a real photograph of Simone Biles).
Figure 1
Most and Least Frequently Identified Images
Textural Analysis
The relationship between texture complexity in images and accuracy in identifying them appears to follow a U-shaped curve (3rd degree polynomial) rather than a linear relationship. Images with medium-high texture ratings (around 7) demonstrated the highest average detection accuracy, 61.0%. Images with very high or low ratings of 8 or 4 showed much lower performance, 18.9% and 38.3% respectively. This suggests a central region where complexity decreases photorealism, and higher and lower regions see less accuracy. Interestingly, this observation, if proven to be true, is consistent with previous research on the uncanny valley phenomenon. Images that have enough detail to appear realistic but are not perfectly photorealistic trigger stronger detection responses than more obvious artificial factors (Kätsyri et al., 2015). That being said, more research is needed to conclusively establish this relationship. Texture ratings themselves are inherently subjective, but assuming validity, the relationship between accuracy is weak, with R² = 0.164. In context, this means that around 16.4% of the variance in percent correct can be explained by texture ratings using a polynomial model. In order to create a confidence interval, bootstrapping with 10,000 iterations and n = 30 was used. The result of this test shows we can be 95% confident that the true R² value for the relationship is between 0.051 and 0.489. This wide range suggests a high level of uncertainty in the model's fit, likely due to the relatively small sample size or the variability in the data itself
Figure 2
Scatter Plot of Accuracy vs. Texture rating, Showing the U-shaped relationship
Texture complexity also exhibited slight positive correlations with the word count of image prompts (R² = 0.15) demonstrating a weak but present relationship between detail in prompts and more textured outputs. The most identified image combinations in the dataset featured moderate texture (rating 6) with concise wording (11-15 words), with 65.7% correct detection. By contrast, images with the highest texture ratings (8) also had among the lowest accuracy rates (21.8% and 16.1%) and notably higher word counts (15-38 words), suggesting that excessive detailing in prompts may increase realism.
Accuracy by Image Category
Accuracy varied slightly across image categories. A one-way ANOVA of accuracy scores, grouped by the six categories, resulted in an F statistic of approximately 0.56 with 5 degrees of freedom between groups and 54 degrees of freedom within groups. This shows that, while there is some variation between groups, there is no statistically significant difference.
Table 3
Mean Accuracy by Image Category
| Category | Accuracy |
|---|---|
| City/Building | 42.60% |
| Object | 45.30% |
| Single Person | 46.80% |
| Nature | 46.90% |
| Multiperson | 51.10% |
| Celebrity/Political | 60.00% |
Confidence Levels
Participants' confidence levels were measured using a 6-point Likert scale at three points during the evaluation: before the first image set (Q1), after feedback on the first set (Q2), and after the second image set (Q3).
Paired-samples t-tests confirm changes in confidence levels. Confidence significantly decreased from Q1 to Q2 (p < 0.001), indicating that the feedback, which revealed participants' actual performance, significantly lowered their initial confidence. A small but non-significant change in confidence was observed from Q2 to Q3 (p = 0.12).
Figure 3
Mean Confidence Levels Across Evaluation Stages
Relationship Between Accuracy and Confidence
A regression analysis was conducted to examine the relationship between overall accuracy (averaged across both image sets) and average confidence (averaged across all three confidence questions). A weak positive correlation was found (R² = 0.067).
Figure 4:
Scatter Plot of Accuracy vs. Confidence
Demographic Factors and Accuracy
Regression analyses were performed to investigate the relationship between demographic variables (age, gender, self-reported "other factors" enhancing ability, and screen time) and accuracy. No statistically significant correlations were found. The examples below highlight this severe lack of correlation.
Age
A very weak, near-zero correlation was found between age and accuracy (R² = 0.006).
Figure 5:
Scatter Plot of Accuracy vs. Age
Gender
No statistically significant difference in accuracy was found between male and female participants (p = 0.295).
Figure 6:
Box and Whisker Plot of Accuracy by Gender
Self-Reported Factors
Skills were self-identified by participants as relevant to the task of identifying images, such as artists, photographers, or people with past experience using generative AI technology. These factors do not show a statistically significant difference from people who did not report such factors (p = 0.567).
Figure 7:
Box and Whisker Plot of Accuracy by Self-Reported Factors
Screen Time
No significant correlation was found between daily screen time and accuracy (R² = 0.000). While most participants reported 0-7 hours/day of screen time with varied accuracy, a few individuals reported much higher screen times (approx. 11-18 hours/day). Despite the higher leverage of these high screen time participants, their accuracy scores were within the standard range observed across the sample, indicating that even extreme reported screen times did not correspond to a change in accuracy within this sample.
Figure 8:
Scatter Plot of Accuracy vs. Screen Time
Neural Network Analysis
To further discover patterns beyond linear relationships, a multilayer perceptron was trained to predict participant accuracy based on demographic factors and confidence metrics. The model demonstrated a validation Mean Squared Error (MSE) of 0.0055, which corresponds to a Root Mean Square Error (RMSE) of 0.0742. In simpler terms, the model’s predictions were off by an average of 7.4% from the true accuracy values participants achieved.
While initial regression analyses found no statistically significant relationships between demographic variables and accuracy (all p > 0.05, R² near 0), the model explained approximately 16% of the variance in accuracy scores (validation R² = 0.1603). This indicates that a moderate amount of variance in accuracy can be explained by non-linear interactions between variables.
Figure 9:
Scatter Plot of Participant Performance and MLP Predicted Values
The results of this study reveal critical insights into human perception of AI-generated images, particularly as diffusion-based models reach unprecedented levels of realism. Participants’ mean accuracy in distinguishing AI-generated images from real photographs was 48.7% before feedback and 48.3% after feedback, with no statistically significant difference between the two image sets. When compared to a random guessing model, t-tests showed no significant difference from chance (p = 0.398 for Set 1, p = 0.876 for Set 2). This finding marks a notable departure from previous research, such as Lu et al. (2023), where humans achieved a low but significant 61.3% accuracy rate in identifying AI-generated images. The decline from 61.3% to effectively random chance (approximately 50%) in this study underscores a major challenge: as AI models like Imagen, FLUX, and others improve, human ability to discern synthetic from real content is diminishing.
This drop to chance-level performance is a groundbreaking observation, as it represents the first documented instance where human accuracy has fallen to random guessing when confronted with cutting-edge diffusion models. Earlier studies, such as Vukojičić et al. (2023), which examined DALL-E 2 outputs, and Lu et al. (2023), which included Stable Diffusion and StyleGAN3, dealt with models that, while advanced for their time, are now considered legacy and have been deprecated or superseded. The photorealism of these newer models likely contributes to this newfound lack of discriminative ability, aligning with the hypothesis that improved model quality would make detection harder over time.
Interestingly, while overall accuracy hovered around chance, individual image analysis showed significant variability (F = 51.41, p < 0.001). This suggests that specific image characteristics, rather than broad categorical differences, are responsible for detection success or failure. While the correlation is not extremely strong, we do have moderate evidence to assume that high image details can contribute to increased image realism. This trend would contrast directly with findings from Miller et al., (2023) who noted that hyperrealism and symmetry in GAN-generated faces often led to the opposite effect: misidentification as real. As Miller’s study exclusively focused on faces and GANs, it’s possible that within a more diverse image dataset, imperfection or exaggerated detail can enhance believability. Variability by image, rather than category (F = 0.56, p > 0.05), indicates that subject matter (e.g., nature vs. city) is less influential than stylistic execution, challenging assumptions from prior research such as Lu et al. (2023), that category-specific visual regions might be important. This finding also demonstrates the importance of effective prompting in the generation of AI images. Adding specific details and textures to prompts can significantly improve photorealism over generic prompts.
Figure 10
Using Prompting to Enhance Realism
The addition of specific details to a prompt is able to significantly increase the photorealism of the resulting image.
The feedback mechanism, intended to improve accuracy, showed no effect (p = 0.627), contrasting with Ali and Qazi (2021), who found modest gains (<10%) in misinformation detection with education. This discrepancy could suggest that a single feedback round is not enough training to improve against rapidly advancing AI. Further research on prolonged exposure or extensive experience with AI image generators may be required to conclusively determine if image identification is truly a learnable skill.
Confidence levels shifted significantly: dropping from 3.6 to 2.2 (p < 0.001) after feedback revealed poor performance, then slightly rebounding to 2.33 (p = 0.12). The weak correlation between accuracy and confidence (R² = 0.067) implies participants are both overconfident in their abilities and struggle to self-assess their skills. This finding further complicates reliance on human ability as AI realism grows.
Demographic factors showed no significant linear correlation with accuracy, contradicting hypotheses that experience (e.g., artists or AI users) might confer an advantage. Interactions with multiple nonlinear relationships add nuance to this finding, however the moderate association of R² = 0.16 remains a weak predictive power. This uniformity across groups reinforces the idea that the challenge lies in the rapidly improving realism of the technology itself, with a lack of clear features that can predict which individuals perform well. This lack of a relationship with demographic factors highlights a new ceiling of difficulty in identifying AI-generated images, and differs from Lu et al. (2023)’s earlier finding of differences between demographics.
The rapid pace of AI development helps contextualize these results. Post-data collection, releases of new image models introduced even higher standards of photorealism
Table 4
Examples of Recent Advancements in AI Generation Models
| Model Name | Type | Noted Advancements | Source |
|---|---|---|---|
| Recraft V3 | Diffusion-based Image Generation | Higher photorealism; text; graphic design tools | (Recraft AI, 2024) |
| FLUX-1.1-pro Ultra/Raw | Diffusion-based Image Generation | Higher photorealism; high resolution generations | (Black Forest Labs, 2024) |
| Imagen3-002 | Diffusion-based Image Generation | Higher photorealism; image editing | (Baldridge et al., 2024) |
| Aurora | Autoregressive Image Generation | Large autoregressive model; celebrity likenesses | (xAI, 2025) |
| Reve 1.0 | Diffusion-based Image Generation | Higher photorealism; prompt understanding | (Reve, 2025) |
| Ideogram 3.0 | Diffusion-based Image Generation | Higher photorealism | (Ideogram, 2025) |
| Gemini 2.0 Flash | LLM-based Image Generation | World knowledge and image understanding | (Kampf & Brichtova, 2025) |
| GPT-4o | LLM-based Image Generation | World knowledge and image understanding | (OpenAI, 2025) |
| Veo 2 | Diffusion-based Video Generation | Stat-of-the-art video generation | (Gupta et al., 2024) |
| Sora | Diffusion-based Video Generation | First major video model release | (Liu et al., 2024) |
| Alibaba Wan | Diffusion-based Video Generation | Natural motion dynamics; open source weights | (Wang et al., 2025) |
| Seaweed | Diffusion-based Video Generation | Realism; visual control; efficiency | (Yang et al., 2025) |
| Seedream 3.0 | Diffusion-based Image Generation | Realism; text rendering; high resolutions; RL training | (ByteDance Seed, 2025) |
A small-scale follow-up evaluation (Appendix F) suggests these advancements maintain the random-level benchmark for human accuracy, though broader testing is needed. This ongoing evolution, alongside unexplored areas like video generation, an area with rapid advancement, and non-photorealistic styles (e.g., paintings, vectors), presents a research landscape with significant options for future inquiry.
Potential limitations influencing the validity of these findings could be examined in several ways. Most notably, this study suffered significantly from convenience sampling. Due to the nature and time frame allowed to complete the project, convenience sampling was relied on as the main method of recruiting participants. This brings up concerns with the generalizability of findings to broader populations. However, it's also important to note that due to the lack of significant correlations found between demographic factors and participant accuracy, it could indicate a higher likelihood of generalization despite the absence of true random sampling. In an initial effort to balance the sample, weighting was used to create a representative sample. The result of this trial was not included in the main body of the paper as it was not notably different from the unweighted sample. This supports the claim of potential generalized ability, but remains the most significant limitation on results. Secondly, certain interpretations, such as the texture ratings of images, are highly subjective and reliant on biased human perceptions. It's likely that these ratings may not accurately represent the true visual complexity of images, and this finding should not be relied on in examinations of this work. Future implementations might implement robust cross confirmation of image ratings, and could look at more reliable indicators of image quality, such as the Fréchet Inception Distance, a common image quality benchmark that compares the distance between feature distributions.
The most notable limitation was the filtering process used in image selection. As the dataset used is not a random sample of generated images, it restricts generalizing findings to all outputs of these models; only to images that are first filtered through a similar selective process.
This study demonstrates a significant new understanding: human ability to distinguish AI-generated images from real photographs has reached a critical threshold. With accuracy falling to random chance (48.7% and 48.3%) when faced with state-of-the-art diffusion models, this marks a significant shift from prior research, where accuracy remained above chance (e.g., 61.3% in Lu et al., 2023), reflecting the increasing photorealism of AI outputs. The lack of improvement post-feedback, coupled with significant variability in individual image identification, suggests that detection is easily avoidable and hinges on subtle stylistic cues like textures and components, rather than model training or demographic factors. As AI content surges (15 billion images, Valyaeva, 2023), this heightens misinformation risks, evident in cases like the 2024 election (Robins-Early, 2024), outpacing current self-reporting regulations on social media (Meta, 2025; X, 2025). Future research should look at video generation, non-photorealistic styles, and create ongoing evaluation necessary to address evolving risks and trust in visual media.
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., … Zheng, X. (2016). TensorFlow: A system for large-scale machine learning (No. arXiv:1605.08695). arXiv. https://doi.org/10.48550/arXiv.1605.08695
Adobe. (2025). AI ethics: Everything you need to know. https://adobe.com/ai/overview/ethics.html
Ali, A., & Qazi, I. A. (2021). Countering misinformation on social media through educational interventions: Evidence from a randomized experiment in Pakistan (No. arXiv:2107.02775). arXiv. https://doi.org/10.48550/arXiv.2107.02775
Altman, S. [@sama]. (2025, March 31). The chatgpt launch 26 months ago was one of the craziest viral moments I'd ever seen, and we added one million users in five days. We added one million users in the last hour [X Post]. X. https://x.com/sama/status/1906771292390666325
Baio, A. (2022, August 30). Exploring 12 million of the 2. 3 billion images used to train Stable Diffusion’s image generator. Waxy.Org. https://waxy.org/2022/08/exploring-12-million-of-the- images-used-to-train-stable-diffusions-image-generator/
Baldridge, J., Bauer, J., Bhutani, M., Brichtova, N., Bunner, A., Castrejon, L., Chan, K., Chen, Y., Dieleman, S., Du, Y., Eaton-Rosen, Z., Fei, H., Freitas, N. de, Gao, Y., Gladchenko, E., Colmenarejo, S. G., Guo, M., Haig, A., Hawkins, W., … Zwols, Y. (2024). Imagen 3 (No. arXiv:2408.07009). arXiv. https://doi.org/10.48550/arXiv.2408.07009
Black Forest Labs. (2024, November 6). Introducing Flux1. 1 [pro] ultra and raw modes. https://blackforestlabs.ai/flux-1-1-ultra/
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324
Brittain, B. (2023, November 30). Artists take new shot at Stability, Midjourney in updated copyright lawsuit. Reuters. https://www.reuters.com/legal/litigation/artists-take-new-shot-stability- midjourney-updated-copyright-lawsuit-2023-11-30/
ByteDance Seed. (2025). Seedream 3.0 Technical Report. arXiv preprint arXiv:2504.11346. https://arxiv.org/abs/2504.11346
Chappell, B. (2025, January 16). LA’s wildfires prompted a rash of fake images. Here’s why. NPR. https://www.npr.org/2025/01/16/nx-s1-5259629/la-wildfires-fake-images
Chui, M., Hazan, E., Roberts, R., Singla, A., Smaje, K., Sukharevsky, A., Yee, L., & Zemmel, R. (2023). The economic potential of generative AI: The next productivity frontier. McKinsey Digital. https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential- of-generative-ai-the-next-productivity-frontier
Creating the unreal: How Nike made its wildest air footwear yet — Nike, inc. (2024, April 11). https://about.nike.com/en/stories/nike-design-athlete-imagined-revolution
Davis, J. (2024). In a digital world with generative ai detection will not be enough. Newhouse Impact Journal, 1, 9–12. https://doi.org/10.14305/jn.29960819.2024.1.1.01
de Graaf, A., & Muros, C. (2025, April 2). Fact check: Fake news on Myanmar, Thailand earthquake – DW – 04/02/2025. Deutsche Welle. https://www.dw.com/en/fact-check-fake-news-on-myanmar- thailand-earthquake/a-72103584
Dhariwal, P., & Nichol, A. (2021). Diffusion models beat GANs on image synthesis. arXiv. https://doi.org/10.48550/ARXIV.2105.05233
DiResta, R., & Goldstein, J. A. (2024). How spammers and scammers leverage ai-generated images on facebook for audience growth (No. arXiv:2403.12838). arXiv. https://doi.org/10.48550/arXiv.2403.12838
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial networks (No. arXiv:1406.2661). arXiv. https://doi.org/10.48550/arXiv.1406.2661
Gupta, A., Razavi, A., Toor, A., Gupta, A., Erhan, D., Shaw, E., Lau, E., Belletti, F., Barth-Maron, G., Shaw, G., Erdogan, H., Sidahmed, H., Nandwani, H., Moraldo, H., Kim, H., Blok, I., Donahue, J., Lezama, J., Mathewson, K., … Chen, Y. (2024). Veo 2. https://deepmind.google/technologies/veo/veo-2/
Holzinger, A., Saranti, A., Angerschmid, A., Finzel, B., Schmid, U., & Mueller, H. (2023). Toward human-level concept learning: Pattern benchmarking for AI algorithms. Patterns, 4(8), 100788. https://doi.org/10.1016/j.patter.2023.100788
Hunter, J. D. (2007). Matplotlib: A 2D graphics environment. Computing in Science & Engineering, 9(3), 90–95. https://doi.org/10.1109/MCSE.2007.55
Ideogram 3.0. (2025, March 26). https://about.ideogram.ai/3.0
Kampf, K., & Brichtova, N. (2025). Experiment with Gemini 2.0 Flash native image generation. Google Developers Blog. https://developers.googleblog.com/en/experiment-with-gemini-20- flash-native-image-generation
Karras, T., Laine, S., & Aila, T. (2019). A style-based generator architecture for generative adversarial networks (No. arXiv:1812.04948). arXiv. https://doi.org/10.48550/arXiv.1812.04948
Kätsyri, J., Förger, K., Mäkäräinen, M., & Takala, T. (2015). A review of empirical evidence on different uncanny valley hypotheses: Support for perceptual mismatch as one road to the valley of eeriness. Frontiers in Psychology, 6. https://doi.org/10.3389/fpsyg.2015.00390
Levine, T. R. (2020). Duped: Truth-default theory and the social science of lying and deception. University of Alabama Press.
Liu, Y., Zhang, K., Li, Y., Yan, Z., Gao, C., Chen, R., Yuan, Z., Huang, Y., Sun, H., Gao, J., He, L., & Sun, L. (2024). Sora: A review on background, technology, limitations, and opportunities of large vision models (No. arXiv:2402.17177). arXiv. https://doi.org/10.48550/arXiv.2402.17177
Lu, Z., Huang, D., Bai, L., Qu, J., Wu, C., Liu, X., & Ouyang, W. (2023). Seeing is not always believing: Benchmarking human and model perception of ai-generated images (No. arXiv:2304.13023). arXiv. https://doi.org/10.48550/arXiv.2304.13023
Lualeperez. (2025). Lualeperez/coursera-introduction-to-deep-learning-with-keras: Coursera: Introduction to deep learning & neural networks with keras. https://github.com/lualeperez/coursera-introduction-to-deep-learning-with-keras
Meta. (2025, April 12). Misinformation Community Standards. Transparency Center. https://transparency.meta.com/policies/community-standards/misinformation/
Miller, E. J., Steward, B. A., Witkower, Z., Sutherland, C. A. M., Krumhuber, E. G., & Dawel, A. (2023). AI hyperrealism: Why ai faces are perceived as more real than human ones. Psychological Science, 34(12), 1390–1403. https://doi.org/10.1177/09567976231207095
OpenAI. (2025). Introducing 4o image generation. OpenAI. https://openai.com/index/introducing- 4o-image-generation
Raghavan, P. (2024, February 23). Gemini image generation got it wrong. We’ll do better. The Keyword; Google. https://blog.google/products/gemini/gemini-image-generation-issue/
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv. https://doi.org/10.48550/ARXIV.2204.06125
Recraft AI. (2024, October 30). Recraft introduces a revolutionary AI model that thinks in design language. https://www.recraft.ai/blog/recraft-introduces-a-revolutionary-ai-model-that-thinks- in-design-language
Reve. (2025, March 24). Halfmoon is Reve Image—And it’s the best image model in the world. https://x.com/reveimage/status/1904211082870456824
Robins-Early, N. (2024, August 26). How did Donald Trump end up posting Taylor Swift deepfakes? The Guardian. https://www.theguardian.com/technology/article/2024/aug/24/trump-taylor- swift-deepfakes-ai
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back- propagating errors. Nature, 323(6088), 533–536. https://doi.org/10.1038/323533a0
Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., & Ganguli, S. (2015). Deep unsupervised learning using nonequilibrium thermodynamics (No. arXiv:1503.03585). arXiv. https://doi.org/10.48550/arXiv.1503.03585
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(56), 1929–1958. http://jmlr.org/papers/v15/srivastava14a.html
TikTok. (2024, April 17). Integrity and authenticity. TikTok Community Guidelines. https://www.tiktok.com/community-guidelines/en/integrity-authenticity
Top websites ranking—Most visited websites in April 2025 | Similarweb. (2025, April 1). [Data Analytics]. Similarweb; Similarweb. https://www.similarweb.com/top-websites/
Valyaeva, A. (2023, August 15). AI image statistics for 2024: How much content was created by Ai. Everypixel Journal. https://journal.everypixel.com/ai-image-statistics
Vukojičić, M., & Veinović, M. (2023, May). Imitation drawing: Can we spot the difference between AI and human generated drawing? https://doi.org/10.15308/Sinteza-2023-23-27
Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.-W., Chen, D., Yu, F., Zhao, H., Yang, J., Zeng, J., Wang, J., Zhang, J., Zhou, J., Wang, J., Chen, J., Zhu, K., Zhao, K., Yan, K., Huang, L., … Liu, Z. (2025). Wan: Open and advanced large-scale video generative models (No. arXiv:2503.20314). arXiv. https://doi.org/10.48550/arXiv.2503.20314
X. (2025, January). Authenticity. X Rules and Policies. https://help.x.com/en/rules-and- policies/authenticity
xAI. (2025, April 6). Grok image generation release | xAI. https://x.ai/news/grok-image-generation- release
Yang, C., Xiao, F., Cheng, F., Chen, H., Guo, H., Wei, M., Zhu, P., Zhao, Q., Lin, S., Zhao, Y., Lin, Z., Qing, Z., Kong, F., Zuo, F., Yan, J., Gui, L., Qi, L., Wang, S., Bi, S., … Yang, J. (2025). Seaweed-7b: Cost-effective training of video generation foundation model [Technical Report]. ByteDance. https://seaweed.video/
AI-Generated Image Dataset
The full dataset of images used in this research is publicly available to download on GitHub and the full set of AI-generated images used in the research is included below.
A diverse range of image types are included in the study, preventing bias towards any one category, which could potentially skew the results. Differentiated image types also assess the ability to detect AI-generated images across different visual domains.
Table A1
Visual Categories used for Image Dataset
| Type | Number of AI images in dataset | Number of real images in dataset |
|---|---|---|
| Nature/Landscape | 5 | 5 |
| Multiperson | 5 | 5 |
| Single person | 5 | 5 |
| Celebrity/political figure | 5 | 5 |
| Architectural/urban landscape | 5 | 5 |
| Object | 5 | 5 |
For use in the evaluation, all images were scaled to a consistent resolution of 1024 × 768 pixels (approximately 0.8 megapixels) in order to control resolution as a potential confounding variable. This choice is also consistent with how most people interact with images online or through social media, where compression and lower resolution images are common.
Model: FLUX-pro (Black Forest labs)
Prompt: “Create a grainy TV quality photograph of Donald Trump talking at a political rally, supporters with trump 2024 signs behind him”
Model: FLUX-pro (Black Forest labs)
Prompt: “A photo of volunteers planting trees in Detroit, wearing green t-shirts.”
Model: Imagen 3 (Google)
Prompt: “A busy dive bar full of people with ambient lighting.”
Model: FLUX-pro (Black Forest labs)
Prompt: “A towering cliffside overlooking a busy highway ocean. A black and white lighthouse stands on the cliff. The sky with muted colors, city lights in the distance”
Model: Imagen 3 (Google)
Prompt: “New York street photograph taken from the road looking up, wide angle lens.”
Model: Imagen 3 (Google)
Prompt: “A closeup photograph of a baton handoff with elite runners, motion blur.”
Model: Recraft V3 (Recraft AI)
Prompt: “Vintage-style photograph of a worker in harsh weather, carrying a basket of fish over their shoulder, standing under heavy rain.”
Model: FLUX-pro (Black Forest labs)
Prompt: “A handwritten quote on a coffee shop blackboard that says ‘Life is a canvas, and you are the brush—paint your own masterpiece.’ Anna Sterling”
Model: FLUX-pro (Black Forest labs)
Prompt: “Nicole Kidman wearing a white dress in front of a media wall, holding an Oscar”
Model: Imagen 3 (Google)
Prompt: “Guys playing soccer in their backyard during the afternoon. Rough, patchy grass with dirt and weeds.”
Model: Imagen 3 (Google).
Prompt: “A woman hanging up a black and white photo in a messy art studio, many other photos hung up on the walls”
Model: Recraft V3 (Recraft AI)
Prompt: “Nighttime shanghai skyline with illuminated skyscrapers reflected in a river.”
Model: Imagen 3 (Google)
Prompt: “A bright Italian coastal town, drone shot”
Model: Recraft V3 (Recraft AI)
Prompt: “A bright mountain landscape reflected in a serene lake. White fluffy clouds in the sky.”
Model: FLUX-1.1-pro (Black Forest labs)
Prompt: “A dense forest with sunlight filtering through the trees.”
Model: FLUX-1.1-pro (Black Forest labs)
Prompt: “A grey, soft focus photo of pale red poppies in a field.”
Model: FLUX-pro (Black Forest labs)
Prompt: “Tom Brady, uniform 12, being lifted up by teammates at a football game.”
Model: Imagen 3 (Google)
Prompt: “A group of friends hiking through a forest trail on a snowy day. Wide angle deep depth of field”
Model: Imagen 3 (Google)
Prompt: “A man walking down the street bird poop on his face, disgusted look”
Model: Imagen 3 (Google)
Prompt: “Create a photo of a lighthouse in Maine”
Model: Imagen 3 (Google)
Prompt: “A close-up of a vintage pocket watch on a rustic wooden table.”
Model: Recraft V3 (Recraft AI)
Prompt: “A pineapple sitting on a rough bench in a tropical setting.”
Model: FLUX-1.1-pro (Black Forest labs)
Prompt: “A long exposure photograph of rocks along a beach with gentle waves.”
Model: FLUX-1.1-pro (Black Forest labs)
Prompt: “A bee collecting pollen from a vibrant flower, its fuzzy body dusted with golden grains.”
Model: FLUX-pro (Black Forest labs)
Prompt: “Elton John performing live on stage with vibrant lighting.”
Model: FLUX-pro (Black Forest labs)
Prompt: “President Joe Biden giving a speech at the United Nations.”
Model: Imagen 3 (Google)
Prompt: “A person jogging along a beach side trail with palm trees during sunrise.”
Model: FLUX-1.1-pro (Black Forest labs)
Prompt: “A Buddhist monk wearing a brown robe, meditating peacefully in a moderately shabby temple with cracked painted concrete, zoomed out to show more of the temple.”
Model: FLUX-pro (Black Forest labs)
Prompt: “A steaming cup of coffee next to an open notebook with a pen.”
Model: FLUX-pro (Black Forest labs)
Prompt: “Artistic photograph of a bicycle parked beside a graffiti-covered wall.”
Real Image Dataset
The full dataset of images used in this research is publicly available to download on GitHub and the full set of real photographs used in the research is included below.
Images were selected based on the same categories described in Appendix A, and processed in a consistent manner. For real images, only photographs with robust EXIF data and a publication date prior to 2021 (the year advanced image generators became mainstream) were selected to use, as this reduces the probability of an AI-generated image being included in the dataset of real photographs.
Details: Photograph of Kendrick Lamar, taken on February 7, 2013. This image was originally published on Wikimedia Commons by Merlijn Hoek. The image is licensed under Creative Commons BY-SA/GFDL.
Details: Photograph of a man standing in front of a bowl and looking towards the left, taken by Clem Onojeghuo on December 3, 2016. This image is free to use and available on Pexels.
Details: Photograph of a group of people enjoying a music concert, taken by Leah Newhouse. The image was uploaded on February 19, 2017, and is free to use. It is available on Pexels.
Details: Photograph of the Statue of Liberty during nighttime, taken by Pierre Blaché on March 22, 2019. This image is free to use and available on Pexels.
Details: Photograph of a cityscape view, taken by Aleksandar Pasaric on February 17, 2016. This image is free to use and available on Pexels.
Details: Photograph of a paintbrush in shallow focus, taken by Daian Gan on April 28, 2010. This image is free to use and available on Pexels.
Details: Close-up photograph of a pomegranate fruit, taken by Roman Odintsov on July 7, 2020. This image is free to use and available on Pexels.
Details: Photograph of a desert during nighttime, taken by Walid Ahmad on January 19, 2018. This image is free to use and available on Pexels.
Details: Photograph of Meryl Streep at the Opening Ceremony of the Tokyo International Film Festival. Taken on October 25, 2016 by Dick Johnson. This image was originally published on Wikimedia Commons. The image is licensed under Creative Commons BY-SA.
Details: Photograph of three persons sitting on stairs talking with each other, taken by Buro Millennial in Leiden, ZH, Netherlands. This image was uploaded on September 21, 2018, and is free to use. It is available on Pexels.
Details: Photograph of a group of women lying on yoga mats under a blue sky, taken by Amin Sujan on June 14, 2018. This image is free to use and available on Pexels.
Details: Photograph of a woman leaning back on a tree trunk using a black DSLR camera during the day, taken by David Bartus on October 7, 2017. This image is free to use and available on Pexels.
Details: Photograph of Saint Basil's Cathedral on New Year's Eve in Moscow, taken by Y Nakanishi on December 31, 2014. The photo was uploaded to Flickr and is licensed with some rights reserved.
Details: Photograph of a white soccer ball, taken by Aphiwat Chuangchoem. This image was uploaded on March 28, 2017, and is free to use. It is available on Pexels.
Details: Photograph of waterfalls in the middle of green trees, taken by Greg Galas on June 12, 2019. This image is free to use and available on Pexels.
Details: Photograph of selective-focus red fruits with snow taken by Nadine Wuchenauer on December 26, 2018. This image is free to use and available on Pexels.
Details: Photograph of President Barack Obama hosting a press conference at the Pentagon in Washington, D.C., on August 4, 2016. Taken by Air Force Tech. Sgt. Brigitte N. Brantley. This image is in the public domain and was uploaded to Flickr.
Details: Photograph of a crowd dancing in a blue-painted enclosure, taken by Maurício Mascaro on March 16, 2018. This image is free to use and available on Pexels. Licensed under Creative Commons.
Details: Photograph of a man in a white T-shirt sitting on a brown rock formation, taken by Tima Miroshnichenko on November 10, 2020. This image is free to use and available on Pexels.
Details: Photograph of a woman wearing a black sleeveless dress holding white headphones during daytime, taken by Tirachard Kumtanom on June 4, 2017. This image is free to use and available on Pexels.
Details: Photograph of a gray spiral building taken from a low angle, taken by Li Jin on October 31, 2019. This image is free to use and available on Pexels.
Details: Close-up photograph of a ukulele, taken on May 25, 2012. This image is free to use under CC0 and is available on Pexels.
Details: Black wooden fence on snow field at a distance of black bare trees taken on February 16, 2012. This image is free to use under the CC0 license and is available on Pexels.
Details: Photograph of a mountain under a cloudy sky, taken by Evgeny Tchebotarev on January 1, 2014. The image is free to use and is available on Pexels.
Details: Photograph of Simone Biles at the 2016 Olympics all-around gold medal podium in Rio de Janeiro. Taken on August 9, 2016, at 18:46 by Agência Brasil Fotografia. This image was originally published on Wikimedia Commons. and is licensed under Creative Commons BY.
Details: Photograph of President Ronald Reagan speaking at a rally for Senator Durenberger on February 8, 1982. Taken by Michael Evans and part of the Ronald Reagan Library collection (C6289-25). This image is in the public domain, sourced from the National Archives.
Details: Photograph of a person standing in front of a brown crate, taken by Clem Onojeghuo on December 3, 2016. This image is free to use and available on Pexels.
Details: Photograph of a man in a black suit riding a bicycle down the street, taken by Andrea Piacquadio. This image was uploaded on January 28, 2019, and is free to use. It is available on Pexels.
Details: Aerial photograph of city buildings near Honolulu, taken by Jess Loiterton on May 21, 2020. This image is free to use and available on Pexels.
Details: Photograph of a pink and white keychain, taken by Marcin Szmigiel on May 23, 2015. This image is free to use and available on Pexels.
F-table of Critical Values
Df1 represents degrees of freedom between groups and Df2 represents degrees of freedom within groups. An F statistic greater than its critical value results in the rejection of the null hypothesis (there is no difference between groups).
Table C1
F-table of Critical Values for α = 0.05
| DF1/DF2 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 15 | 20 | 24 | 30 | 40 | 60 | 120 | ∞ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2 | 161 | 199 | 215 | 224 | 230 | 233 | 236 | 238 | 240 | 241 | 243 | 245 | 248 | 249 | 250 | 251 | 252 | 253 | 254 |
| 3 | 18.5 | 19 | 19.1 | 19.2 | 19.3 | 19.3 | 19.3 | 19.3 | 19.3 | 19.4 | 19.4 | 19.4 | 19.4 | 19.4 | 19.4 | 19.4 | 19.4 | 19.4 | 19.5 |
| 4 | 10.1 | 9.55 | 9.28 | 9.12 | 9.01 | 8.94 | 8.89 | 8.85 | 8.81 | 8.79 | 8.74 | 8.7 | 8.66 | 8.64 | 8.62 | 8.59 | 8.57 | 8.55 | 8.53 |
| 5 | 7.71 | 6.94 | 6.59 | 6.39 | 6.26 | 6.16 | 6.09 | 6.04 | 6 | 5.96 | 5.91 | 5.86 | 5.8 | 5.77 | 5.75 | 5.72 | 5.69 | 5.66 | 5.63 |
| 6 | 6.61 | 5.79 | 5.41 | 5.19 | 5.05 | 4.95 | 4.88 | 4.82 | 4.77 | 4.74 | 4.68 | 4.62 | 4.56 | 4.53 | 4.5 | 4.46 | 4.43 | 4.4 | 4.37 |
| 7 | 5.99 | 5.14 | 4.76 | 4.53 | 4.39 | 4.28 | 4.21 | 4.15 | 4.1 | 4.06 | 4 | 3.94 | 3.87 | 3.84 | 3.81 | 3.77 | 3.74 | 3.7 | 3.67 |
| 8 | 5.59 | 4.74 | 4.35 | 4.12 | 3.97 | 3.87 | 3.79 | 3.73 | 3.68 | 3.64 | 3.57 | 3.51 | 3.44 | 3.41 | 3.38 | 3.34 | 3.3 | 3.27 | 3.23 |
| 9 | 5.32 | 4.46 | 4.07 | 3.84 | 3.69 | 3.58 | 3.5 | 3.44 | 3.39 | 3.35 | 3.28 | 3.22 | 3.15 | 3.12 | 3.08 | 3.04 | 3.01 | 2.97 | 2.93 |
| 10 | 5.12 | 4.26 | 3.86 | 3.63 | 3.48 | 3.37 | 3.29 | 3.23 | 3.18 | 3.14 | 3.07 | 3.01 | 2.94 | 2.9 | 2.86 | 2.83 | 2.79 | 2.75 | 2.71 |
| 15 | 4.6 | 3.74 | 3.34 | 3.11 | 2.96 | 2.85 | 2.76 | 2.7 | 2.65 | 2.6 | 2.53 | 2.46 | 2.39 | 2.35 | 2.31 | 2.27 | 2.22 | 2.18 | 2.13 |
| 20 | 4.38 | 3.52 | 3.13 | 2.9 | 2.74 | 2.63 | 2.54 | 2.48 | 2.42 | 2.38 | 2.31 | 2.23 | 2.16 | 2.11 | 2.07 | 2.03 | 1.98 | 1.93 | 1.88 |
| 30 | 4.18 | 3.33 | 2.93 | 2.7 | 2.55 | 2.43 | 2.35 | 2.28 | 2.22 | 2.18 | 2.1 | 2.03 | 1.94 | 1.9 | 1.85 | 1.81 | 1.75 | 1.7 | 1.64 |
| 40 | 4.08 | 3.23 | 2.84 | 2.61 | 2.45 | 2.34 | 2.25 | 2.18 | 2.12 | 2.08 | 2 | 1.92 | 1.84 | 1.79 | 1.74 | 1.69 | 1.64 | 1.58 | 1.51 |
| 60 | 4 | 3.15 | 2.76 | 2.53 | 2.37 | 2.25 | 2.17 | 2.1 | 2.04 | 1.99 | 1.92 | 1.84 | 1.75 | 1.7 | 1.65 | 1.59 | 1.53 | 1.47 | 1.39 |
| 120 | 3.92 | 3.07 | 2.68 | 2.45 | 2.29 | 2.18 | 2.09 | 2.02 | 1.96 | 1.91 | 1.83 | 1.75 | 1.66 | 1.61 | 1.55 | 1.5 | 1.43 | 1.35 | 1.25 |
| ∞ | 3.84 | 3 | 2.6 | 2.37 | 2.21 | 2.1 | 2.01 | 1.94 | 1.88 | 1.83 | 1.75 | 1.67 | 1.57 | 1.52 | 1.46 | 1.39 | 1.32 | 1.22 | 1 |
MLP Training Script
The following Python script implements a multilayer perceptron for predicting participant accuracy based on demographic and confidence variables. This script was based on an original implementation as a part of the Introduction to Deep Learning and Neural Networks with Keras, a course offered by IBM (lualeperez, 2025). It utilizes the TensorFlow and Keras frameworks (Abadi et al., 2016) and modeled after the original neural network design (Rumelhart et al., 1986). The feature importance analysis is adapted from Breiman's variable importance measures (Breiman, 2001).
The script employs standard practices (Srivastava et al., 2014), uses Matplotlib (Hunter, 2007).
# Import statements set up necessary tools and libraries. These tools allow the implementation of the model as described above.
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.metrics import mean_squared_error, r2_score
import tensorflow as tf
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Dropout
from tensorflow.keras.callbacks import EarlyStopping
# A random seed is used to set up both numpy and tensorflow. Because the seed is fixed, this ensures reproducibility by future researchers, and similar results should be produced each time the script is run.
np.random.seed(42)
tf.random.set_seed(42)
# To analyze a dataset, data is imported from a file set as “demographics.csv” which has the nine input demographic factors as the first nine columns. Factors were converted to quantitative values before analysis.
print("Loading data from demographics.csv...")
data = pd.read_csv('demographics.csv')
print(f"Dataset shape: {data.shape}")
# Ensures that there were not any mistakes in the data formatting. (see previous comment)
print("\nData types of each column:")
print(data.dtypes)
# Check again for mistakes
print("\nMissing values in each column:")
print(data.isnull().sum())
# the rest of the script
X = data.iloc[:, :9] # First 9 columns are the demo. Factors for each participant.
y = data.iloc[:, 9] # The 10th column is each participant’s accuracy on the evaluation.
# Splits data into training and validation sets. A validation set ensures the model is actually generalizable in the real world. A fairly large 20% validation set is used in this case.
X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)
# Standardize the input features using z scores!
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_val_scaled = scaler.transform(X_val)
# Create model is created here
print("Creating the model...")
model = Sequential([
Dense(16, activation='relu', input_shape=(9,)), # This is the first hidden layer, which has 16 neurons and uses ReLU activation.
Dropout(0.2), # Randomly turns off some neurons, which prevents overfitting, a concern as we only have 280 participants as training examples.
Dense(8, activation='relu'), # Hidden layer 2, with 8 neurons.
Dense(1, activation='sigmoid') # Output neuron, which compresses the values into a probability between 1 and 0 using a sigmoid function. Because probabilities are values on a scale from 0 to 1, the sigmoid simply normalizes the model’s prediction into an easier to understand metric.
])
# This model uses an adam optimizer for training, which is standard practice. Instead of optimizing a loss function, this model is designed to minimize the difference between predictions and actual participant’s accuracy.
model.compile(
optimizer='adam',
loss='mean_squared_error', # MSE for regression
metrics=['mae']
)
# Print function
model.summary()
# This sets up early stopping to avoid overfitting. The script will wait 30 training cycles without improvement and keep track of the validation loss.
early_stopping = EarlyStopping(
monitor='val_loss',
patience=30,
restore_best_weights=True
)
print("\nTraining the model...")
history = model.fit(
X_train_scaled, y_train,
epochs=200,
batch_size=16, # Small batch size as this is a small dataset. If this research is reproduced with an extremely large sample size increasing this from 16 might help speed up training.
validation_data=(X_val_scaled, y_val),
callbacks=[early_stopping],
verbose=1
)
# Evaluate the model’s performance using the validation dataset.
print("\nEvaluating model performance...")
y_pred_val = model.predict(X_val_scaled).flatten()
val_mse = mean_squared_error(y_val, y_pred_val)
val_rmse = np.sqrt(val_mse)
val_r2 = r2_score(y_val, y_pred_val)
print(f"Validation MSE: {val_mse:.4f}")
print(f"Validation RMSE: {val_rmse:.4f}")
print(f"Validation R²: {val_r2:.4f}")
# Make predictions using the entire dataset, in order to compare to actual values.
X_all_scaled = scaler.transform(X)
y_pred = model.predict(X_all_scaled).flatten()
# Create a plot of the training history to observe how well the model learns.
plt.figure(figsize=(12, 5))
plt.subplot(1, 2, 1)
plt.plot(history.history['loss'], label='Training Loss')
plt.plot(history.history['val_loss'], label='Validation Loss')
plt.title('Model Loss During Training')
plt.xlabel('Epoch')
plt.ylabel('Loss (MSE)')
plt.legend()
plt.grid(True)
plt.subplot(1, 2, 2)
plt.plot(history.history['mae'], label='Training MAE')
plt.plot(history.history['val_mae'], label='Validation MAE')
plt.title('Model MAE During Training')
plt.xlabel('Epoch')
plt.ylabel('Mean Absolute Error')
plt.legend()
plt.grid(True)
plt.tight_layout()
plt.savefig('train.png')
plt.show()
# Create a regression analysis of actual values vs values predicted by the model.
plt.figure(figsize=(10, 6))
plt.scatter(y, y_pred, alpha=0.5)
plt.xlabel('Actual Values')
plt.ylabel('Predicted Values')
plt.title('MLP Prediction Performance')
plt.grid(True)
plt.savefig('regression.png')
plt.show()
print("\nAnalysis completed!")Survey Materials
Data was collected using Google Forms. The survey and questions are publicly available on Forms and the questions are included below. Participants were presented with the following information before consenting.
IntroductionI am a student at Morse High School, and I am in a class that requires me to complete a research study. My project is to investigate how well individuals can distinguish between AI-generated images and real photographs. You have been selected as a possible participant because you are part of community outreach efforts. There are no specific exclusionary criteria; anyone aged 18 or older is welcome to participate. Please read this form before agreeing to participate in this study. Purpose of the Study The purpose of the study is to examine how effectively people can differentiate between AI-generated images and real photographs, and to understand whether receiving feedback improves their ability and confidence in subsequent attempts. I intend to present the findings of this study in a written academic paper submitted to the College Board, aligning with their terms. I may submit the results for publication in a student research journal. If you participate in this study, you will be asked to complete an online survey consisting of two sets of images. In the first set, you will view 30 images and classify each one as either real or AI-generated. After completing the first set, you will receive feedback on your responses. Then, you will evaluate a second set of 30 images in the same manner. Throughout the study, you will also be asked to rate your confidence levels using three scale questions ranging from 1 to 6. The entire study will take approximately 5–10 minutes to complete. Risks and Benefits The study involves minimal risks. There are no foreseeable psychological or physical risks, but as with any research, there may be unknown risks that are currently unforeseeable. There are no direct benefits, gifts, or payments for participating in this study. Your participation will contribute to research that may help improve understanding of how people perceive AI-generated content. Confidentiality This study is confidential, and I will not be retaining or reporting any information about your identity. Findings from the study will not include any information that would make it possible to identify you, and results will only be reported in aggregate. The records of this study will be kept strictly secure. Data will be stored on a password-protected computer accessible only by me. No audio or video recordings will be made. By agreeing to participate, you consent to the inclusion of the study findings in any reports or publications. Questions and Concerns You have the right to ask questions about this research study and to have those questions answered before or after your participation. If you have any further questions about the study, feel free to contact me, Declan Wright, at declan.wright@rsu1.org. A summary of the results of the study can be sent to you upon request. Survey Questions Email Certification I have read the above terms and consent to voluntary participation in this research. Background Information What year were you born? (e.g., 1986) What is your gender? Male Female Other: What is your current education level? Have you ever encountered an image online that you suspected was AI-generated? Yes No In which state do you currently reside? If you are outside the United States, please select "International." Approximately how many hours per day do you spend on your smartphone? (Average daily screen time for the last full week, rounded to the nearest hour) Are there any other factors that might be significant to your ability to discern AI-generated and real photographs? (e.g., photographer, artist, familiarity with AI tools, tech industry employee) Please share them now. Before answering any questions, how confident are you in your ability to tell the difference between real and AI-generated images? Image Classification Set (Round 1) (Participants were shown 30 images, one at a time) 11-40. For each image presented (Image 1 through Image 30): Do you think this image is: AI-generated Real photograph Feedback Section (Participants were shown the following instructions and the correct answers for images 11-40) Answers to the first set of images: Review your mistakes and analyze image features as necessary. DO NOT change answers to the previous section under any circumstances. (Correct classifications for the 30 images were displayed here) After reviewing the correct answers from the first set of images, please rate your confidence in distinguishing between AI-generated images and authentic photographs. Image Classification Set (Round 2) (Participants were shown the second set of 30 images, one at a time) 42-71. For each image presented (Image 31 through Image 60): Do you think this image is: AI-generated Real photograph Closing Please rate your confidence in distinguishing between AI-generated images and authentic photographs after completing the second set of images. How helpful was the feedback you received after completing the first set of images in improving your ability to distinguish between AI-generated images and authentic photographs? Type the word "blue" before continuing. (End of Survey) |
Small-Scale Evaluation with Newer Image Models
Following the completion of the initial round of data collection for this study, several newer and more advanced AI image generation models were released by developers. These models represent strong improvements in photorealism and detail compared to those used in the main evaluation dataset. To assess whether the main finding, that human detection accuracy had fallen to chance levels, remains true against these state-of-the-art models, a small-scale follow-up evaluation was conducted in early 2025.
Table F1
Models Used in Image Generation
| Model | Developer | Release Date | Images Generated |
|---|---|---|---|
| Imagen3-002 | Google DeepMind | February 5, 2025 | 40% |
| Recraft V3* | Recraft AI | October 30, 2024 | 37% |
| FLUX 1.1 [pro] Raw | Black Forest Labs | November 6, 2024 | 13% |
| Aurora | xAI Corp. | December 9, 2024 | 10% |
*Includes Photoreal, Hard Flash, and Black & White finetuned versions of the model.
The evaluation format mirrored the main study, presenting participants with a total of 60 real photographs and synthetic images from the newer models above. For simplicity, images were not split into multiple feedback groups, and reduced demographic information was collected. A small group of participants (n = 15) completed this follow-up.
The results from this small-scale evaluation indicated that participant accuracy remained near chance levels (mean accuracy of 48.5%). This finding aligns closely with the original observations of the main study (48.7% and 48.3% accuracy). Sample t-tests result in a p value of 0.90. While a small sample size does add some uncertainty to this finding, the effect size from the sample also shows us something interesting: if a difference were to truly exist in the broader population, we would require a sample size of around 1.6 million people to confidently observe that difference, which highlights how minuscule even a true change would be. In other words, the critical threshold of human ability to detect AI-generated images at random levels persists, and is potentially solidified with the very latest advancements in diffusion model technology available at the time of the follow-up. It is important to note that the field continues to evolve at an extremely fast pace, with even newer models and architectures (such as image generation powered by large language models) being released regularly even after this follow-up was conducted.
Interestingly, while the mean accuracy remained stable, there is an observed change in the variance of participant scores occurred between the main study and this follow-up. The overall standard deviation (of all 60 images together) decreased from 7.5% in the original research to 5.4% in the follow-u. To formally evaluate the difference in variances, Levene’s Test was used. For each observation in each group, Levene's test calculates the deviation from its group's median, then performs an ANOVA on these values. This resulted in a W statistic of 2.58.
As shown in Figure F1, this observed W statistic falls below the critical values required for statistical significance at commonly used alpha levels (α = 0.05 and α = 0.10). Therefore, we do not have enough evidence to reject the null hypothesis that the variances are equal; the observed difference is not statistically significant.
Figure F1
Levene’s Test F Distribution With Selected Values
It is notable that the observed W statistic of 2.58 approaches the critical value for α = 0.10. While failing to meet the threshold, this proximity might suggest a potential trend towards reduced variability in detection accuracy when participants face the very latest models. A lower variance might suggest that images created by the next generation of image models are more consistently challenging, due to the tighter association around the chance-level mean. This hypothesis is speculative, but may be worth additional research, particularly given the limitation of the very small sample size in this follow-up study. Such a small sample is highly susceptible to random chance, and a larger sample would be required to evaluate this possibility.
Overall, this follow-up builds further evidence to support the main study: human accuracy in detecting state-of-the-art AI-generated images is currently indistinguishable from random chance. The observed decrease in score variance and its proximity to significance thresholds deserves mention, but does not suggest any new understanding within itself, other than that the field's rapid advancement requires continued evaluation of human limits against AI advancement.