Blur of Reality: Evaluating Human Detection of Cutting-Edge AI-Generated Images

Declan Wright, May 17, 2025









Introduction

History

ChatGPT’s release in November 2022 marked a rush of interest in AI development, capturing widespread public and academic attention. It sparked what many consider the next age of internet innovation, with massive economic opportunity for those involved. What started out as a research project evolved into a commercial product embraced by millions worldwide, surpassing initial projections of growth. It is now the seventh most used website globally, ranking only behind the likes of Google, Instagram and Facebook (Similarweb, 2025). Technologies powering ChatGPT and other tools are collectively known as generative artificial intelligence (AI). These models can create novel content that replicates human writing, speech, and visual media with realism. A foundational innovation leading to the generative models used today was conceptualized in the 1980s following simultaneous independent discoveries of the multilayer perceptron, or MLP (Rumelhart et al., 1986). This discovery is what allows AI models to learn patterns. They work by updating internal parameters based on observed examples. MLPs are loosely modeled after the human nervous system, but are primarily mathematical and statistical models designed for nonlinear pattern recognition. It is notable that while humans typically require one or two examples to learn something (Holzinger et al., 2023), training powerful neural networks requires massive amounts of data. Stable Diffusion, a popular open source image generator, was trained on 2.3 billion images (Baio, 2022).

In 2014, the first landmark AI technology was released for image generation. Known as Generative Adversarial Networks, or GANs, these models were a novel framework proposed by Goodfellow and colleagues at the University of Montreal (2014). GANs are trained through an adversarial process between two neural networks: a generator that creates images, and a discriminator that tries to distinguish between real and generated images. This setup creates a competitive dynamic where the generator improves at creating more realistic images to fool the discriminator, while the discriminator becomes better at detecting fakes. As an early innovation, GANs often struggled with training stability and model collapse. They also have other limitations. StyleGAN, a model developed by Nvidia in 2018, can only generate images of faces. Users cannot specify attributes like age, ethnicity, or expression, which restricts its flexibility and makes it less adaptable compared to newer models which can generate a wide variety of images (Karras et al., 2019).  


Diffusion Models

Diffusion models, introduced more recently, have largely superseded GANs as the primary technology for image generation. The mathematical foundation for diffusion was conceptualized by researchers at UC Berkeley (Sohl-Dickstein et al., 2015). Adapted to image training, diffusion models are trained by adding fixed amounts of random noise to training images. In photography, noise refers to random pixels that disrupt the structure of an image. Diffusion models learn to reverse this noise addition process, enabling them to start with completely random pixels and progressively remove noise by predicting and subtracting what the noise should look like at each step of the generation process. This approach offers several advantages over GANs. Diffusion models are more stable during training and don't suffer from the same model collapse issues. They also can generate a wider variety of image concepts with better consistency (Dhariwal & Nichol, 2021). Additionally, the introduction of the CLIP encoder by OpenAI enabled diffusion models with even greater flexibility, allowing text prompting, more precise control over the generation process, and features like image editing (Ramesh et al., 2022). This enhanced controllability has led to explosive adoption. An estimated 34 million AI images were generated daily as of 2023, with approximately 15 billion AI images created by that time (Valyaeva, 2023). This number has grown exponentially since then. Following the release of a new image generation model in ChatGPT, the platform saw over a million new users sign up for the product in a single hour (Altman, 2025). While Instagram hosts approximately 50 billion images accumulated over more than a decade, AI-generated imagery has now likely exceeded that figure. Traditional stock photography platforms have been completely outpaced. Shutterstock's entire library of 386 million images represents less than a month's worth of current AI image production (Valyaeva, 2023).


Outlook and Concerns

The newfound popularity of image models has created significant risks. The findings presented in this study aim to evaluate the level of concern that may be appropriate for the widespread distribution and use of advanced image models. One of the most widely cited concerns with AI models is the potential for misinformation. During the 2024 US presidential election, prominent figures including Candidates in the 2024 US presidential election repeatedly posted AI-generated photos of other candidates and celebrities such as Taylor Swift, a newfound concern with voter influence (Robins-Early, 2024). Extensive wildfires in Los Angeles were almost immediately followed by a flood of AI-generated images online, depicting iconic landmarks like the Hollywood sign on fire, when they were not at risk from the fires (Chappell, 2025). More recently, similar engagement baiting content generated by AI depicting false scenes of earthquakes in Myanmar received millions of cumulative views online (de Graaf & Muros, 2025). In the past, creating manipulated images required extensive technical skill. Today, AI models allow images to be created with little skill, taking seconds to process. A 2024 preprint conducted by Stanford researchers found a significant wave of engagement-baiting content generated by AI on Facebook. This content receives hundreds of millions or engagements, with the majority of users not realizing they were viewing artificial content. Perhaps more concerningly, the Facebook algorithm appears to be actively promoting these posts, increasingly presenting them to a wider range of audiences (DiResta & Goldstein, 2024). This raises significant concerns about the ethical implications of algorithmic amplification, particularly how it enables the spread of misinformation and its potential to erode public trust in media. In publicly released community guidelines, Facebook itself, as well as other social media platforms including TikTok and X, all have similar policies on AI content: placing the burden on users to self-report AI usage (Meta, 2025; TikTok, 2024; X, 2025). Consequently, regulating AI misinformation will likely require a comprehensive approach. A meta-review of AI visual content concluded that effective regulation demands both watermarked credentials on photos and the deployment of AI detection systems working in tandem, as neither method alone represents a foolproof approach (Davis, 2024). Another promising approach when it comes to combating misinformation is education. Even basic knowledge of misinformation can help people improve their ability to identify it, despite gains remaining under 10% improvement (Ali & Qazi, 2021). This idea is consistent with historical trends as well. Even in ancient societies, human populations adapt to new forms of deception extremely quickly (Levine, 2020).

Image generators face other concerns. The usage of billions of images in training diffusion models has led to widespread outrage among artists and photographers whose work may have been used in model training. Highly publicized lawsuits, such as the 2023 challenge reported by Brittain (2023), have brought further awareness to this concern. While the issue of copyright is yet to be decided in the courts, some companies have taken a proactive approach to training their models. Adobe, the company behind popular image editing software such as Photoshop, stated in a 2024 release that its Firefly models were trained by paying artists for rights to use their images, and allowing users to have their images opted out of AI training (Adobe, 2025). Another issue has been the introduction of racial biases. Models trained on images that lean heavily towards a specific demographic group, or are tuned to produce output based on race, can have notable consequences. Following the release of Google’s flagship AI product Gemini, the company was forced to release a public apology and stop image generation features on the platform, after the model generated racially inappropriate images of certain historical groups (Raghavan, 2024). This incident not only sparked widespread criticism from the public but also raised concerns within the tech industry about the lack of rigorous testing for bias in AI models. It pressured Google to re-evaluate its training datasets and  slow the rollout of similar features in subsequent projects. This feature was restored with the release of Imagen 3, Google’s latest model, which featured extensive de-biasing within training data and in tuning before being released (Baldridge et al., 2024).

While the risks associated with AI image generation are notable, the technology also offers tangible benefits. In the case of image generators, previous academic analysis of their impacts has mostly focused on economic benefits, with generative AI predicted to add trillions of dollars per year to the global economy (Chui et al., 2023). Investment in AI technologies has seen incredible interest, even during a period of slow general economic growth (Chui et al., 2023).

 For creative professionals and designers, the use of AI is democratizing. For those without the previous knowledge and ability to create visuals, AI tools can bring ideas into the world. This concept also brings promise to marketing and prototyping. Nike worked with a group of elite athletes who were able to visualize prototypes and concepts using AI tools. This significantly accelerated the development process and allowed the Nike design team to bring athletes’ visualizations to reality (Creating the unreal, 2024).


Identifying AI-Generated Images

Previous research on an individual’s ability to differentiate between AI-generated images and real photographs is severely lacking. A Serbian study (Vukojičić et al., 2023) conducted by student researchers was the first directly presenting research on this topic in early 2023. The authors compared how well people could tell the difference between images created using DALL-E 2 and hand drawn sketches. This research found that performance varied greatly depending on the difficulty of the task, and whether images were presented in groups or in pairs. While it provides a good basis for understanding the topic, its sample was limited to a small group or students, and the method of presenting image pairs introduces bias that could have influenced the ability of participants to discern AI-generated drawings. More recent research involving collaborators across several institutions from across China, Hong Kong, and Sydney, evaluated human performance at discerning AI-generated images, and compared it to a model trained on the Fake2M dataset (Lu et al., 2023). In this research, humans were able to identify AI images correctly 61.3% of the time. The images in that dataset were created using Stable Diffusion, StyleGAN3, and DeepFloyd IF, a multi-process model. This does present some issues about the generalizability of the findings, as models like GANs and some forms of auto-regression are considered antiquated technologies. This directly leads to the development of a research question: “How well can individuals distinguish between photographs created using leading-edge generative diffusion models as compared to naturally created photographs, and how does providing feedback on their initial performance impact their confidence and ability to improve in subsequent attempts?” This guiding question aims to address these gaps in previous research, by using leading-edge diffusion models such as Imagen 3 and the FLUX family of models. Compared to generators used in previous research, such as Stable Diffusion and Dall-E 2, these leading-edge models bring significant improvements in image quality and photorealism, representing an area that remains understudied. This evaluation also addresses two other significant gaps in research. Firstly, as previous research has found positive effects on people’s ability to identify misinformation, a feedback mechanism will be implemented, allowing participants to observe their mistakes and potentially learn from them. Secondly, participants’ confidence will be measured at three points during the research process. Unlike previous studies which took the ability of participants at face value, this provides an additional dimension for interpreting the results.




Methods 

This study aims to determine whether individuals can accurately identify AI-generated images. To achieve this, a dataset comprising photorealistic images generated by AI alongside real photographs was compiled, and a human survey using the images was conducted. The following sections explain the specifics of the data collection process, the human evaluation, and the techniques employed to analyze the findings.

Image Dataset

A total of 60 images were compiled to create the dataset used in this research. The images consist of 30 real photographs and 30 synthetic images, generated using advanced diffusion models. 

Table 1

Models Used in Image Generation

ModelDeveloperRelease DateImages Generated (%)
Imagen 3Google DeepMindMay 14, 202431%
FLUX.1 ProBlack Forest LabsAugust 1, 202435%
FLUX 1.1 ProBlack Forest LabsOctober 2, 202417%
Recraft V3Recraft AIOctober 30, 202417%


After a diverse range of images were generated, the highest quality images were selected for use as a part of this research. This selective process does not result in random samples of generated images, but was chosen to mirror real-world environments where AI-generated images are optimized for “believability” and potential misinformation. Examples of images that were not used include images with obvious artifacts (a hand with ten fingers), or a lack of fine textures or general realism. To put this in perspective, for every image used in the dataset, an average of 4.83 additional images were generated and not used.

To prevent participants from making direct comparisons that could bias their decisions, the real and synthetic image datasets do not contain image pairs (e.g., there is not a synthetic image of a black cat paired with a real photograph of a black cat). Authentic photographs were sourced from reputable, free-use websites including Wikimedia Commons, Pixabay, and Flickr. All real images were published before 2021, as this predates the availability of modern AI image generators. To avoid bias in selecting images, the photographs are distributed evenly across six different categories with five images each. These categories are designed to represent a diverse array of image depictions across different visual regions, to ensure that there is not a strong leading or bias in the image data set (Lu et al., 2023). The use of any graphic, suggestive, or other sensitive images was avoided in order to ensure participants’ safety and comfort with the evaluation. Full details regarding each image, including file names, sources, AI model details (including prompts), visual categories, and resolutions, are provided in Appendix A.

In order to evaluate perceptual qualities of images, a texture and complexity rating scale from one to ten was created. Images with unique complexities or strong textures were given higher scores closer to ten, whereas images with low levels of texture or specific details were given ratings closer to one. These ratings were assigned by manually coding the images.

Evaluation Design

Images are divided randomly into two groups, each in separate survey sections. Google Forms was used to build and host the evaluation. Participants were recruited through community outreach, with a total of  n = 280 individuals completing the evaluation. After the first set of images was evaluated, feedback with the correct answers to these initial photos was given to participants. The second set of images was evaluated under the same method as the first. Basic demographic questions were also asked for the purpose of analyzing differences across such demographic groups, specifically looking at gender, age, education level, location of residence, screen time, and a self-reported question about any factors that might enhance a participant’s ability (for example, artists, or people with experience using AI). Participants’ confidence levels were measured using three quantitative questions ranging from 1 to 6. This scale was chosen as it pushes participants to a certain result instead of allowing them to reply with a neutral answer, which is a common issue in survey design. The first question is placed before the initial image set, the second after reviewing the feedback from the first set, and the final question after the completion of the second image set. This approach is intended to show changes in participants’ confidence, which potentially offers an additional way to interpret ability beyond accuracy. Informed consent was included as a part of the evaluation. The complete instructions provided to participants, details about the feedback content, and the exact wording of the demographic and Likert scale questions can be found in Appendix B.


Statistical Analysis

Following the collection of data, the results were analyzed using Google Sheets. Data analysis was conducted in several ways. Initially, descriptive statistics were employed. Raw data from the evaluation was converted into numerical values. Correct responses were graded with a 1, and incorrect responses received a 0. Each participant’s overall accuracy was found by averaging these results and data was described using standard deviations.

Subsequently, a “random guessing” model was created. To build the random model, 30 random integers (either 0 or 1) were generated, and the mean of these integers was taken to represent a simulated participant’s score. This process was repeated 280 times to create a normal distribution of simulated participants with the same sample size as the real evaluation.

Statistical analysis methods, such as sample t-tests, were used to determine if the difference between the two image sets and the random model were statistically significant. A t-test is a statistical procedure used to determine whether the mean difference between two observations is convincing evidence to reject a null hypothesis. In this case, t-tests were used to compare participant accuracy before and after feedback, and were also used to compare confidence between individual questions. 

Additionally, Analysis of Variance (ANOVA) was used to compare means across multiple categories simultaneously. Unlike t-tests which can only compare two groups, ANOVA allows for a similar comparison of performance across more than two groups. This method helped identify if there were statistically significant differences in accuracy among image categories and at the individual image level.

Finally, a regression analysis was used to determine the relationships between scores and confidence, and scores and demographic information. Statistical significance was modeled using the coefficient of determination , which measures the percentage of variance in the data explained by the model. In addition, a multilayer perceptron (MLP) neural network was trained to model participant accuracy based on demographic factors and confidence ratings. The MLP consisted of two hidden layers (16 and 8 neurons) with rectified linear unit (Leaky ReLU) activation functions. See Appendix D for the full training script with comments explaining its precise implementation. This approach allowed for discovery of non-linear relationships that might not be apparent through standard regression analysis.

This statistical approach effectively addressed the research question by directly measuring participants’ ability to distinguish between AI-generated and real images, the effectiveness of feedback in enhancing detection skills, and changes in confidence levels and their correlation with performance.




Results

 The following section presents research findings, focusing on participant accuracy, changes in confidence levels, and the relationships between these variables and demographic factors.

Overall Accuracy

 Participants' ability to correctly identify AI-generated images was assessed across two sets of images. The mean accuracy for image set one, before feedback, was 48.7%.  For image set two, which participants completed after receiving feedback, the mean was 48.3%. A paired-sample t-test resulted in a p value of 0.627, which is not convincing evidence for statistically significant difference between the accuracy rates of the two image sets.

 Furthermore, participants' performance was compared to a random guessing model. Sample t-tests indicated that accuracy for both image sets was not statistically different from chance.  This suggests that, on average, participants were unable to reliably distinguish between real and AI-generated images, even after receiving feedback.


Table 2
Descriptive Statistics for Accuracy

Image SetMean Accuracy (%)Standard Deviation (%)p-value (vs. random model)
Set 1 (before feedback)48.7%10.1%0.398
Set 2 (after feedback)48.3%9.6%0.876

Note:  A standard significance level of α = 0.05 was used to reject the null hypothesis.

Individual Image Analysis

While overall accuracy was near chance levels, a one-way ANOVA revealed highly  statistically significant differences in correct identification rates between individual images. This resulted in an F statistic of approximately 51.41 with 59 degrees of freedom between groups and 16,770 degrees of freedom within groups. Accuracy rates ranged from a low 12.1% (for an AI-generated image of a woman painter) to a high of 88.6% (for a real photograph of Simone Biles).

Figure 1
Most and Least Frequently Identified Images

Most and least frequently identified images020406080100Image NamePercent CorrectA11: 12.1%A11A25: 16%A25A28: 17.4%A28A9: 18.9%A9B5: 21.1%B5A27: 75.6%A27B17: 77.6%B17B23: 78.6%B23A19: 81.5%A19B25: 88.6%B25
Figure or table reproduced from the original research paper


Textural Analysis

The relationship between texture complexity in images and accuracy in identifying them appears to follow a U-shaped curve (3rd degree polynomial) rather than a linear relationship. Images with medium-high texture ratings (around 7) demonstrated the highest average detection accuracy, 61.0%. Images with very high or low ratings of 8 or 4 showed much lower performance, 18.9% and 38.3% respectively. This suggests a central region where complexity decreases photorealism, and higher and lower regions see less accuracy. Interestingly, this observation, if proven to be true, is consistent with previous research on the uncanny valley phenomenon. Images that have enough detail to appear realistic but are not perfectly photorealistic trigger stronger detection responses than more obvious artificial factors (Kätsyri et al., 2015). That being said, more research is needed to conclusively establish this relationship. Texture ratings themselves are inherently subjective, but assuming validity, the relationship between accuracy is weak, with = 0.164. In context, this means that around 16.4% of the variance in percent correct can be explained by texture ratings using a polynomial model. In order to create a confidence interval, bootstrapping with 10,000 iterations and n = 30 was used. The result of this test shows we can be 95% confident that the true value for the relationship is between 0.051 and 0.489. This wide range suggests a high level of uncertainty in the model's fit, likely due to the relatively small sample size or the variability in the data itself


Figure 2
Scatter Plot of Accuracy vs. Texture rating, Showing the U-shaped relationship

Accuracy versus texture rating34567890%50%100%Texture RatingPercent Correct6, 58.21435, 53.57145, 60.35715, 30.71435, 34.64294, 756, 64.28576, 58.21434, 80.71437, 28.21435, 87.85714, 65.71435, 73.21435, 756, 74.28574, 26.07145, 22.14295, 46.07146, 44.28575, 43.92867, 31.78575, 53.57146, 21.07148, 78.21435, 11.42864, 61.07146, 37.58, 83.92867, 53.92867, 42.1429


Texture complexity also exhibited slight positive correlations with the word count of image prompts ( = 0.15) demonstrating a weak but present relationship between detail in prompts and more textured outputs. The most identified image combinations in the dataset featured moderate texture (rating 6) with concise wording (11-15 words), with 65.7% correct detection. By contrast, images with the highest texture ratings (8) also had among the lowest accuracy rates (21.8% and 16.1%) and notably higher word counts (15-38 words), suggesting that excessive detailing in prompts may increase realism.


Accuracy by Image Category

Accuracy varied slightly across image categories. A one-way ANOVA of accuracy scores, grouped by the six categories, resulted in an F statistic of approximately 0.56 with 5 degrees of freedom between groups and 54 degrees of freedom within groups. This shows that, while there is some variation between groups, there is no statistically significant difference.

Table 3
Mean Accuracy by Image Category

CategoryAccuracy
City/Building42.60%
Object45.30%
Single Person46.80%
Nature46.90%
Multiperson51.10%
Celebrity/Political60.00%


Confidence Levels

Participants' confidence levels were measured using a 6-point Likert scale at three points during the evaluation: before the first image set (Q1), after feedback on the first set (Q2), and after the second image set (Q3). 

Paired-samples t-tests confirm changes in confidence levels. Confidence significantly decreased from Q1 to Q2 (p < 0.001), indicating that the feedback, which revealed participants' actual performance, significantly lowered their initial confidence. A small but non-significant change in confidence was observed from Q2 to Q3 (p = 0.12).

Figure 3
Mean Confidence Levels Across Evaluation Stages

Mean confidence across evaluation stages01234Confidence LevelConfidence Q13.60Confidence Q22.20Confidence Q32.30


Relationship Between Accuracy and Confidence

A regression analysis was conducted to examine the relationship between overall accuracy (averaged across both image sets) and average confidence (averaged across all three confidence questions). A weak positive correlation was found ( = 0.067).

Figure 4:
Scatter Plot of Accuracy vs. Confidence

Accuracy versus average confidence1234560%20%40%60%80%100%Average ConfidencePercent Correct3, 404.33333, 604, 46.66672, 602, 604, 43.33332.66667, 56.66673.33333, 503.66667, 63.33332, 36.66673.66667, 63.33332.33333, 66.66671.66667, 43.33332, 46.66672.33333, 401, 36.66673.33333, 46.66671, 53.33333.33333, 36.66672.33333, 53.33333.33333, 603.33333, 46.66671.66667, 704, 46.66674, 36.66673.66667, 401, 36.66672, 53.33333, 46.66672.66667, 603, 401.66667, 43.33332.66667, 36.66672.33333, 43.33334, 33.33334, 53.33332.33333, 36.66672, 46.66671.66667, 403, 56.66672.33333, 53.33332, 502.33333, 502.66667, 602.33333, 53.33332.33333, 503.33333, 56.66673.33333, 33.33331.66667, 402.66667, 56.66672.33333, 56.66673.66667, 46.66672.66667, 46.66672.33333, 43.33333.66667, 46.66674, 63.33332.33333, 602.33333, 56.66672.33333, 53.33334, 36.66672.33333, 56.66671.66667, 43.33333, 43.33332.66667, 702.33333, 36.66671.33333, 503.66667, 502.33333, 33.33332.33333, 402, 502.33333, 504.33333, 46.66673.33333, 602.66667, 56.66671.66667, 53.33333, 33.33332.33333, 502, 56.66673, 53.33333.66667, 502.33333, 43.33333.66667, 63.33333, 603.33333, 66.66672.33333, 602, 401.33333, 502.33333, 43.33333, 56.66674, 46.66673.66667, 46.66671.33333, 46.66672.33333, 53.33332.33333, 46.66673.33333, 56.66674.33333, 56.66672.33333, 43.33333.66667, 46.66674, 603.33333, 46.66673.33333, 66.66676, 66.66673.66667, 504.33333, 53.33332.66667, 66.66673, 46.66672.66667, 63.33332.66667, 43.33331.33333, 63.33332.33333, 503, 703.33333, 46.66672.33333, 63.33331.66667, 602.33333, 404, 43.33332.33333, 603, 502.66667, 402, 56.66673, 36.66673.33333, 36.66672.66667, 502.66667, 404.33333, 46.66673.33333, 46.66672.33333, 502.33333, 46.66672.33333, 43.33332, 402.66667, 402.33333, 53.33332.33333, 402.33333, 53.33333, 43.33332, 46.66672, 402.33333, 43.33334.33333, 56.66673.33333, 33.33331.33333, 402.66667, 53.33333, 501.66667, 501, 401, 53.33331.66667, 46.66672.66667, 43.33332.66667, 53.33331.66667, 36.66672.66667, 53.33331.33333, 46.66673, 43.33332.33333, 43.33333, 33.33333, 43.33333, 302.33333, 36.66671.33333, 43.33332, 36.66672, 56.66671, 36.66673.33333, 63.33334, 503.33333, 56.66674.33333, 403.33333, 402.33333, 43.33333.33333, 46.66672.33333, 53.33333.33333, 401.33333, 43.33333.33333, 46.66671.66667, 33.33332.33333, 403.33333, 46.66672.66667, 301.33333, 302.33333, 36.66672, 301.66667, 401.33333, 33.33332, 33.33331.33333, 504, 63.33332.33333, 43.33332.66667, 46.66672, 36.66671.33333, 33.33333.33333, 26.66674.33333, 46.66674.33333, 56.66673, 46.66671.33333, 503, 501.33333, 43.33332.66667, 53.33331, 36.66671, 46.66673.66667, 504, 66.66672.33333, 503.66667, 501.33333, 404, 43.33331.66667, 43.33331.66667, 402.33333, 403, 36.66672.66667, 43.33332, 63.33332, 56.66671.33333, 33.33333.33333, 53.33332.33333, 63.33331.33333, 46.66671.66667, 46.66672.33333, 46.66671.33333, 53.33332, 503, 603.33333, 53.33333.33333, 46.66672, 601.66667, 404, 56.66672.33333, 36.66672.66667, 46.66672.66667, 43.33332, 503.33333, 602, 403.33333, 502.33333, 53.33334.33333, 46.66672.66667, 46.66673, 56.66674.33333, 43.33333, 602.66667, 46.66672, 504, 603.33333, 56.66672.33333, 43.33331, 56.66675.33333, 56.66673.33333, 43.33332, 303, 23.33335, 83.33333, 46.66673.33333, 404, 46.66672, 43.33333, 53.33332.66667, 502.66667, 56.66673.33333, 63.33332.33333, 26.66672, 53.33331.66667, 53.33332.66667, 33.33331.66667, 33.33331, 43.33332, 43.33333, 402.66667, 43.33332, 43.33332.33333, 36.66673.33333, 63.33333.33333, 56.66673.33333, 56.66673.33333, 66.66673.66667, 502, 53.33331.66667, 56.66675, 46.66674.33333, 56.66674.33333, 56.66674.33333, 70


Demographic Factors and Accuracy

Regression analyses were performed to investigate the relationship between demographic variables (age, gender, self-reported "other factors" enhancing ability, and screen time) and accuracy. No statistically significant correlations were found. The examples below highlight this severe lack of correlation.

Age

A very weak, near-zero correlation was found between age and accuracy ( = 0.006). 

Figure 5:
Scatter Plot of Accuracy vs. Age

Accuracy versus age2040608010010%20%30%40%50%60%70%80%Age (years)Percent Correct18, 36.666718, 53.333318, 6018, 53.333318, 5018, 56.666718, 43.333318, 6018, 4058, 66.666718, 46.666718, 33.333387, 56.666747, 6027, 4075, 33.333322, 46.666747, 5049, 6036, 6028, 46.666726, 36.666719, 63.333347, 46.666736, 56.666776, 53.333318, 63.333341, 46.666719, 5025, 56.666732, 46.666774, 33.333342, 46.666757, 4053, 4045, 53.333325, 23.333344, 4043, 46.666726, 4043, 53.333362, 63.333350, 56.666718, 56.666724, 4042, 63.333335, 56.666759, 46.666766, 5069, 26.666745, 53.333318, 5077, 5058, 6040, 4018, 43.333319, 6019, 4018, 5018, 46.666718, 3018, 53.333318, 63.333328, 53.333377, 4061, 56.666718, 5042, 5043, 43.333344, 53.333333, 56.666742, 46.666718, 4049, 36.666749, 56.666755, 5043, 3018, 63.333353, 46.666720, 36.666729, 53.333321, 6068, 63.333318, 33.333343, 33.333344, 43.333351, 26.666745, 56.666718, 53.333352, 7024, 46.666737, 63.333339, 43.333318, 46.666745, 6036, 5072, 53.333335, 7035, 53.333322, 6024, 6030, 56.666722, 66.666730, 63.333355, 6032, 56.666739, 5022, 5022, 56.666758, 56.666718, 6066, 53.333319, 33.333346, 43.333335, 43.333332, 63.333318, 6022, 5035, 43.333320, 5025, 56.666745, 53.333324, 43.333317, 6032, 5044, 36.666718, 46.666749, 5017, 26.666743, 43.333316, 5032, 36.666758, 5041, 6044, 3059, 56.666752, 33.333349, 46.666719, 56.666762, 56.666764, 33.333350, 53.333357, 6044, 63.333348, 56.666765, 53.333367, 4051, 43.333336, 56.666732, 36.666755, 3047, 53.333359, 36.666777, 26.666750, 5054, 46.666736, 56.666724, 5044, 6029, 56.666718, 5048, 6035, 53.333352, 63.333318, 43.333340, 56.666760, 36.666752, 33.333363, 4077, 53.333334, 36.666733, 4051, 46.666753, 4039, 36.666746, 3041, 36.666721, 46.666770, 53.333350, 3066, 36.666740, 53.333380, 5056, 5032, 63.333375, 53.333358, 5046, 46.666765, 46.666795, 53.333357, 43.333351, 46.666770, 4075, 66.666735, 53.333350, 5042, 56.666781, 36.666774, 6050, 6026, 53.333367, 6049, 6045, 46.666749, 33.333337, 36.666757, 46.666754, 3049, 56.666755, 46.666750, 56.666730, 36.666738, 43.333349, 63.333328, 4045, 46.666747, 5054, 5056, 4053, 5045, 56.666738, 6062, 33.333319, 43.333364, 6025, 3049, 4053, 53.333355, 66.666761, 66.666766, 3061, 6048, 33.333352, 56.666746, 6056, 6027, 5018, 53.333318, 53.333330, 4068, 43.333366, 46.666759, 5067, 43.333382, 46.666732, 36.666765, 4031, 33.333382, 33.333339, 4059, 63.333360, 53.333350, 36.666755, 43.333357, 43.333338, 6018, 4043, 6029, 36.666736, 5028, 5039, 36.666777, 33.333365, 43.333367, 4026, 4034, 36.666748, 5035, 5017, 43.333318, 43.333318, 6017, 46.666724, 43.333346, 33.333328, 6024, 56.666718, 56.666718, 66.666718, 66.6667


Gender

No statistically significant difference in accuracy was found between male and female participants (p = 0.295).

Figure 6:
Box and Whisker Plot of Accuracy by Gender

Accuracy by gender30%40%50%60%70%Percent CorrectMaleFemale


Self-Reported Factors

Skills were self-identified by participants as relevant to the task of identifying images, such as artists, photographers, or people with past experience using generative AI technology. These factors do not show a statistically significant difference from people who did not report such factors (p = 0.567).

Figure 7:
Box and Whisker Plot of Accuracy by Self-Reported Factors

Accuracy by self-reported factors30%40%50%60%70%Percent CorrectOther Factors = YesOther Factors = No


Screen Time

No significant correlation was found between daily screen time and accuracy ( = 0.000). While most participants reported 0-7 hours/day of screen time with varied accuracy, a few individuals reported much higher screen times (approx. 11-18 hours/day). Despite the higher leverage of these high screen time participants, their accuracy scores were within the standard range observed across the sample, indicating that even extreme reported screen times did not correspond to a change in accuracy within this sample.

Figure 8:
Scatter Plot of Accuracy vs. Screen Time

Accuracy versus screen time05101520%30%40%50%60%70%80%Screen Time (hours/day)Percent Correct4, 36.66676, 53.33332, 603, 53.33333, 501, 56.66674, 43.33334, 604, 402, 66.66675, 46.66675, 33.33332, 56.66674, 604, 402, 33.33338, 46.66676, 502, 607, 603, 46.66672, 36.66672, 63.33335, 46.66674, 56.666718, 53.33331, 63.33332, 46.66674, 504, 56.66673, 46.66671, 33.33335, 46.66671, 402, 404, 53.33332, 23.33334, 405, 46.66674, 401, 53.33332, 63.33332, 56.66676, 56.66674, 403, 63.33338, 56.66673, 46.66675, 503, 26.66673, 53.33335, 502, 504, 606, 405, 43.33333, 605, 405, 507, 46.66674, 303, 53.33331, 63.33332, 53.33333, 401, 56.66674, 503, 503, 43.33335, 53.33333, 56.66675, 46.66677, 404, 36.66671, 56.66671, 504, 304, 63.33332, 46.66675, 36.66674, 53.33333, 601, 63.33333, 33.33333, 33.33331, 43.33334, 26.66673, 56.66674, 53.33332, 703, 46.66676, 63.333312, 43.33335, 46.66672, 605, 501, 53.33332, 704, 53.33336, 605, 605, 56.66674, 66.66675, 63.33332, 604, 56.66676, 503, 502, 56.66671, 56.66675, 603, 53.33333, 33.33333, 43.33332, 43.33332, 63.33332, 605, 503, 43.33332, 502, 56.66675, 53.33338, 43.33333, 604, 501, 36.66675, 46.66672, 506, 26.66673, 43.33333, 504, 36.66673, 503, 603, 302, 56.66672, 33.33333, 46.66673, 56.66673, 56.66675, 33.33332, 53.33331, 600, 63.33330, 56.66672, 53.33333, 402, 43.33333, 56.66673, 36.66671, 304, 53.33332, 36.66672, 26.66673, 502, 46.66674, 56.66672, 503, 604, 56.66675, 504, 604, 53.33333, 63.33338, 43.33336, 56.66674, 36.66675, 33.333318, 402, 53.33332, 36.66673, 401, 46.66675, 404, 36.66672, 303, 36.66673, 46.66673, 53.33333, 303, 36.66672, 53.33333, 503, 508, 63.33333, 53.33331, 502, 46.66673, 46.66672, 53.33332, 43.33332, 46.66674, 402, 66.66678, 53.33335, 504, 56.66670, 36.66676, 605, 607, 53.33333, 603, 605, 46.66673, 33.33333, 36.66672, 46.66673, 303, 56.66672, 46.66672, 56.66672, 36.66673, 43.33334, 63.33333, 404, 46.66672, 502, 505, 407, 504, 56.66676, 602, 33.33336, 43.33334, 603, 302, 404, 53.33332, 66.66672, 66.66674, 303, 602, 33.33332, 56.66676, 602, 605, 506, 53.33335, 53.33331, 404, 43.33331, 46.66672, 504, 43.33331, 46.66677, 36.66671, 403, 33.33333, 33.33333, 402, 63.33332, 53.33332, 36.66676, 43.33332, 43.33332, 605, 401, 601, 36.66671, 501, 502, 36.66672, 33.33331, 43.33333, 402, 402, 36.66670, 501, 505, 43.33332, 43.33332, 603, 46.66675, 43.33333, 33.33333, 6011, 56.66674, 56.66674, 66.66673, 66.6667



Neural Network Analysis

To further discover patterns beyond linear relationships, a multilayer perceptron was trained to predict participant accuracy based on demographic factors and confidence metrics.  The model demonstrated a validation Mean Squared Error (MSE) of 0.0055, which corresponds to a Root Mean Square Error (RMSE) of 0.0742. In simpler terms, the model’s predictions were off by an average of 7.4% from the true accuracy values participants achieved.

While initial regression analyses found no statistically significant relationships between demographic variables and accuracy (all p > 0.05, near 0), the model explained approximately 16% of the variance in accuracy scores (validation = 0.1603). This indicates that a moderate amount of variance in accuracy can be explained by non-linear interactions between variables. 

Figure 9:
Scatter Plot of Participant Performance and MLP Predicted Values

Participant performance and saved MLP predictions304050607035%40%45%50%55%60%65%Actual ValuesPredicted Values38.3333, 47.283156.6667, 53.679253.3333, 55.142256.6667, 50.260655, 53.468350, 55.283250, 59.881355, 55.283951.6667, 55.270751.6667, 47.391855, 52.497450, 49.507450, 47.389853.3333, 46.420740, 43.880435, 41.054546.6667, 48.316651.6667, 48.155448.3333, 46.554856.6667, 48.937353.3333, 50.371841.6667, 47.579366.6667, 55.21546.6667, 49.317646.6667, 48.495746.6667, 52.690350, 48.775350, 49.279848.3333, 53.345358.3333, 49.713343.3333, 43.038138.3333, 48.043841.6667, 43.78841.6667, 47.863136.6667, 44.765853.3333, 48.773930, 45.190843.3333, 43.631443.3333, 47.30748.3333, 44.074253.3333, 47.360956.6667, 49.434453.3333, 48.640558.3333, 51.525146.6667, 43.85356.6667, 45.375956.6667, 52.2240, 47.878245, 45.654641.6667, 46.945755, 46.6148.3333, 52.56948.3333, 51.272751.6667, 45.932743.3333, 45.735353.3333, 52.653760, 48.779348.3333, 49.144951.6667, 47.928741.6667, 48.259743.3333, 48.962948.3333, 48.605253.3333, 49.933261.6667, 49.473938.3333, 45.502953.3333, 44.045850, 56.106741.6667, 44.06541.6667, 46.556551.6667, 49.939153.3333, 48.308246.6667, 52.863850, 49.842246.6667, 48.945255, 49.654941.6667, 47.297140, 47.491260, 47.825250, 42.129143.3333, 52.532648.3333, 48.466361.6667, 55.176561.6667, 47.497550, 52.635946.6667, 48.581941.6667, 43.379238.3333, 40.925950, 46.735355, 53.233758.3333, 52.799746.6667, 50.891155, 49.44948.3333, 53.142146.6667, 51.041158.3333, 50.410953.3333, 53.181548.3333, 45.933658.3333, 49.613156.6667, 54.524553.3333, 52.438963.3333, 54.624161.6667, 56.870458.3333, 53.383658.3333, 52.964963.3333, 49.706551.6667, 49.138756.6667, 51.566646.6667, 49.737760, 46.858253.3333, 49.926465, 54.978450, 47.500848.3333, 45.51651.6667, 45.827241.6667, 43.713953.3333, 51.668160, 54.278550, 53.731841.6667, 46.640353.3333, 47.261346.6667, 49.377745, 47.094746.6667, 48.244850, 45.648148.3333, 53.18441.6667, 49.08948.3333, 46.248348.3333, 47.222735, 45.22641.6667, 45.101645, 48.699145, 47.437845, 44.514656.6667, 48.637236.6667, 43.829751.6667, 45.541636.6667, 47.788745, 45.046656.6667, 49.914845, 47.604936.6667, 41.526553.3333, 48.652755, 47.294656.6667, 49.867148.3333, 47.384253.3333, 47.913643.3333, 45.662343.3333, 51.522255, 47.379336.6667, 43.687741.6667, 47.772850, 43.778640, 43.716735, 44.486141.6667, 46.596245, 48.496943.3333, 50.878643.3333, 46.71751.6667, 43.857846.6667, 41.257653.3333, 54.596248.3333, 46.645758.3333, 48.185356.6667, 52.232550, 48.329948.3333, 51.712838.3333, 43.987338.3333, 45.872543.3333, 51.101453.3333, 46.911838.3333, 43.061141.6667, 43.877946.6667, 49.615536.6667, 45.654538.3333, 48.03838.3333, 46.176333.3333, 52.618738.3333, 45.131245, 46.8330, 41.671538.3333, 46.854343.3333, 43.203141.6667, 39.897950, 48.064663.3333, 54.923848.3333, 46.703548.3333, 45.048841.6667, 45.418640, 45.08340, 37.653645, 49.539651.6667, 50.747543.3333, 47.973658.3333, 46.330651.6667, 52.629646.6667, 44.583955, 47.363236.6667, 46.754253.3333, 44.362255, 51.769760, 53.815355, 47.918955, 47.82443.3333, 49.291938.3333, 46.681840, 44.525443.3333, 40.752735, 40.62246.6667, 45.879345, 46.630960, 43.094946.6667, 49.20638.3333, 44.870758.3333, 52.081651.6667, 48.618846.6667, 44.528348.3333, 45.978348.3333, 46.179646.6667, 48.987750, 50.177458.3333, 50.087756.6667, 52.655640, 47.939251.6667, 48.029650, 44.29743.3333, 51.853438.3333, 40.243650, 45.897255, 48.574258.3333, 47.612145, 48.645550, 47.32541.6667, 43.829355, 48.648753.3333, 52.953453.3333, 48.821153.3333, 51.99248.3333, 53.799956.6667, 42.232943.3333, 50.703146.6667, 46.997953.3333, 49.531653.3333, 49.206743.3333, 45.442451.6667, 46.231946.6667, 54.622241.6667, 49.225531.6667, 43.912928.3333, 41.649261.6667, 56.172755, 46.459346.6667, 49.539941.6667, 46.318343.3333, 46.734548.3333, 43.86455, 48.595748.3333, 46.593461.6667, 52.102331.6667, 39.59351.6667, 46.429351.6667, 45.247535, 47.689833.3333, 47.295843.3333, 45.665541.6667, 45.070840, 46.979740, 48.061846.6667, 46.270643.3333, 45.273853.3333, 51.476450, 53.425458.3333, 53.425456.6667, 51.167346.6667, 52.71243.3333, 45.377558.3333, 48.218851.6667, 54.379956.6667, 54.455561.6667, 54.455568.3333, 56.5513




Analysis and Discussion

The results of this study reveal critical insights into human perception of AI-generated images, particularly as diffusion-based models reach unprecedented levels of realism. Participants’ mean accuracy in distinguishing AI-generated images from real photographs was 48.7% before feedback and 48.3% after feedback, with no statistically significant difference between the two image sets. When compared to a random guessing model, t-tests showed no significant difference from chance (p = 0.398 for Set 1, p = 0.876 for Set 2). This finding marks a notable departure from previous research, such as Lu et al. (2023), where humans achieved a low but significant 61.3% accuracy rate in identifying AI-generated images. The decline from 61.3% to effectively random chance (approximately 50%) in this study underscores a major challenge: as AI models like Imagen, FLUX, and others improve, human ability to discern synthetic from real content is diminishing.

This drop to chance-level performance is a groundbreaking observation, as it represents the first documented instance where human accuracy has fallen to random guessing when confronted with cutting-edge diffusion models. Earlier studies, such as Vukojičić et al. (2023), which examined DALL-E 2 outputs, and Lu et al. (2023), which included Stable Diffusion and StyleGAN3, dealt with models that, while advanced for their time, are now considered legacy and have been deprecated or superseded. The photorealism of these newer models likely contributes to this newfound lack of discriminative ability, aligning with the hypothesis that improved model quality would make detection harder over time.

Interestingly, while overall accuracy hovered around chance, individual image analysis showed significant variability (F = 51.41, p < 0.001). This suggests that specific image characteristics, rather than broad categorical differences, are responsible for detection success or failure. While the correlation is not extremely strong, we do have moderate evidence to assume that high image details can contribute to increased image realism. This trend would contrast directly with findings from Miller et al., (2023) who noted that hyperrealism and symmetry in GAN-generated faces often led to the opposite effect: misidentification as real. As Miller’s study exclusively focused on faces and GANs, it’s possible that within a more diverse image dataset, imperfection or exaggerated detail can enhance believability. Variability by image, rather than category (F = 0.56, p > 0.05), indicates that subject matter (e.g., nature vs. city) is less influential than stylistic execution, challenging assumptions from prior research such as Lu et al. (2023), that category-specific visual regions might be important. This finding also demonstrates the importance of effective prompting in the generation of AI images. Adding specific details and textures to prompts can significantly improve photorealism over generic prompts. 

Figure 10
Using Prompting to Enhance Realism

The addition of specific details to a prompt is able to significantly increase the photorealism of the resulting image.


The feedback mechanism, intended to improve accuracy, showed no effect (p = 0.627), contrasting with Ali and Qazi (2021), who found modest gains (<10%) in misinformation detection with education. This discrepancy could suggest that a single feedback round is not enough training to improve against rapidly advancing AI. Further research on prolonged exposure or extensive experience with AI image generators may be required to conclusively determine if image identification is truly a learnable skill.

Confidence levels shifted significantly: dropping from 3.6 to 2.2 (p < 0.001) after feedback revealed poor performance, then slightly rebounding to 2.33 (p = 0.12). The weak correlation between accuracy and confidence ( = 0.067) implies participants are both overconfident in their abilities and struggle to self-assess their skills. This finding further complicates reliance on human ability as AI realism grows.

Demographic factors showed no significant linear correlation with accuracy, contradicting hypotheses that experience (e.g., artists or AI users) might confer an advantage. Interactions with multiple nonlinear relationships add nuance to this finding, however the moderate association of = 0.16 remains a weak predictive power. This uniformity across groups reinforces the idea that the challenge lies in the rapidly improving realism of the technology itself, with a lack of clear features that can predict which individuals perform well. This lack of a relationship with demographic factors highlights a new ceiling of difficulty in identifying AI-generated images, and differs from Lu et al. (2023)’s earlier finding of differences between demographics. 

The rapid pace of AI development helps contextualize these results. Post-data collection, releases of new image models introduced even higher standards of photorealism

Table 4
Examples of Recent Advancements in AI Generation Models

Model NameTypeNoted AdvancementsSource
Recraft V3Diffusion-based Image GenerationHigher photorealism; text; graphic design tools(Recraft AI, 2024)
FLUX-1.1-pro Ultra/RawDiffusion-based Image GenerationHigher photorealism; high resolution generations(Black Forest Labs, 2024)
Imagen3-002Diffusion-based Image GenerationHigher photorealism; image editing(Baldridge et al., 2024)
AuroraAutoregressive Image GenerationLarge autoregressive model; celebrity likenesses(xAI, 2025)
Reve 1.0Diffusion-based Image GenerationHigher photorealism; prompt understanding(Reve, 2025)
Ideogram 3.0Diffusion-based Image GenerationHigher photorealism(Ideogram, 2025)
Gemini 2.0 FlashLLM-based Image GenerationWorld knowledge and image understanding(Kampf & Brichtova, 2025)
GPT-4oLLM-based Image GenerationWorld knowledge and image understanding(OpenAI, 2025)
Veo 2Diffusion-based Video GenerationStat-of-the-art video generation(Gupta et al., 2024)
SoraDiffusion-based Video GenerationFirst major video model release(Liu et al., 2024)
Alibaba WanDiffusion-based Video GenerationNatural motion dynamics; open source weights(Wang et al., 2025)
SeaweedDiffusion-based Video GenerationRealism; visual control; efficiency(Yang et al., 2025)
Seedream 3.0Diffusion-based Image GenerationRealism; text rendering; high resolutions; RL training(ByteDance Seed, 2025)



A small-scale follow-up evaluation (Appendix F) suggests these advancements maintain the random-level benchmark for human accuracy, though broader testing is needed. This ongoing evolution, alongside unexplored areas like video generation, an area with rapid advancement, and non-photorealistic styles (e.g., paintings, vectors), presents a research landscape with significant options for future inquiry.

Potential limitations influencing the validity of these findings could be examined in several ways. Most notably, this study suffered significantly from convenience sampling. Due to the nature and time frame allowed to complete the project, convenience sampling was relied on as the main method of recruiting participants. This brings up concerns with the generalizability of findings to broader populations. However, it's also important to note that due to the lack of significant correlations found between demographic factors and participant accuracy, it could indicate a higher likelihood of generalization despite the absence of true random sampling. In an initial effort to balance the sample, weighting was used to create a representative sample. The result of this trial was not included in the main body of the paper as it was not notably different from the unweighted sample. This supports the claim of potential generalized ability, but remains the most significant limitation on results. Secondly, certain interpretations, such as the texture ratings of images, are highly subjective and reliant on biased human perceptions. It's likely that these ratings may not accurately represent the true visual complexity of images, and this finding should not be relied on in examinations of this work. Future implementations might implement robust cross confirmation of image ratings, and could look at more reliable indicators of image quality, such as the Fréchet Inception Distance, a common image quality benchmark that compares the distance between feature distributions. 

The most notable limitation was the filtering process used in image selection. As the dataset used is not a random sample of generated images, it restricts generalizing findings to all outputs of these models; only to images that are first filtered through a similar selective process.





Conclusion

This study demonstrates a significant new understanding: human ability to distinguish AI-generated images from real photographs has reached a critical threshold. With accuracy falling to random chance (48.7% and 48.3%) when faced with state-of-the-art diffusion models, this marks a significant shift from prior research, where accuracy remained above chance (e.g., 61.3% in Lu et al., 2023), reflecting the increasing photorealism of AI outputs. The lack of improvement post-feedback, coupled with significant variability in individual image identification, suggests that detection is easily avoidable and hinges on subtle stylistic cues like textures and components, rather than model training or demographic factors. As AI content surges (15 billion images, Valyaeva, 2023), this heightens misinformation risks, evident in cases like the 2024 election (Robins-Early, 2024), outpacing current self-reporting regulations on social media (Meta, 2025; X, 2025). Future research should look at video generation, non-photorealistic styles, and create ongoing evaluation necessary to address evolving risks and trust in visual media.





References

Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., … Zheng, X. (2016). TensorFlow: A system for large-scale machine learning (No. arXiv:1605.08695). arXiv. https://doi.org/10.48550/arXiv.1605.08695 

Adobe. (2025). AI ethics: Everything you need to know. https://adobe.com/ai/overview/ethics.html

Ali, A., & Qazi, I. A. (2021). Countering misinformation on social media through educational interventions: Evidence from a randomized experiment in Pakistan (No. arXiv:2107.02775). arXiv. https://doi.org/10.48550/arXiv.2107.02775

Altman, S. [@sama]. (2025, March 31). The chatgpt launch 26 months ago was one of the craziest viral moments I'd ever seen, and we added one million users in five days. We added one million users in the last hour [X Post]. X. https://x.com/sama/status/1906771292390666325

Baio, A. (2022, August 30). Exploring 12 million of the 2. 3 billion images used to train Stable Diffusion’s image generator. Waxy.Org. https://waxy.org/2022/08/exploring-12-million-of-the- images-used-to-train-stable-diffusions-image-generator/

Baldridge, J., Bauer, J., Bhutani, M., Brichtova, N., Bunner, A., Castrejon, L., Chan, K., Chen, Y., Dieleman, S., Du, Y., Eaton-Rosen, Z., Fei, H., Freitas, N. de, Gao, Y., Gladchenko, E., Colmenarejo, S. G., Guo, M., Haig, A., Hawkins, W., … Zwols, Y. (2024). Imagen 3 (No. arXiv:2408.07009). arXiv. https://doi.org/10.48550/arXiv.2408.07009

Black Forest Labs. (2024, November 6). Introducing Flux1. 1 [pro] ultra and raw modes. https://blackforestlabs.ai/flux-1-1-ultra/

Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324

Brittain, B. (2023, November 30). Artists take new shot at Stability, Midjourney in updated copyright lawsuit. Reuters. https://www.reuters.com/legal/litigation/artists-take-new-shot-stability- midjourney-updated-copyright-lawsuit-2023-11-30/

ByteDance Seed. (2025). Seedream 3.0 Technical Report. arXiv preprint arXiv:2504.11346. https://arxiv.org/abs/2504.11346

Chappell, B. (2025, January 16). LA’s wildfires prompted a rash of fake images. Here’s why. NPR. https://www.npr.org/2025/01/16/nx-s1-5259629/la-wildfires-fake-images

Chui, M., Hazan, E., Roberts, R., Singla, A., Smaje, K., Sukharevsky, A., Yee, L., & Zemmel, R. (2023). The economic potential of generative AI: The next productivity frontier. McKinsey Digital. https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential- of-generative-ai-the-next-productivity-frontier

Creating the unreal: How Nike made its wildest air footwear yet — Nike, inc. (2024, April 11). https://about.nike.com/en/stories/nike-design-athlete-imagined-revolution

Davis, J. (2024). In a digital world with generative ai detection will not be enough. Newhouse Impact Journal, 1, 9–12. https://doi.org/10.14305/jn.29960819.2024.1.1.01

de Graaf, A., & Muros, C. (2025, April 2). Fact check: Fake news on Myanmar, Thailand earthquake – DW – 04/02/2025. Deutsche Welle. https://www.dw.com/en/fact-check-fake-news-on-myanmar- thailand-earthquake/a-72103584

Dhariwal, P., & Nichol, A. (2021). Diffusion models beat GANs on image synthesis. arXiv. https://doi.org/10.48550/ARXIV.2105.05233

DiResta, R., & Goldstein, J. A. (2024). How spammers and scammers leverage ai-generated images on facebook for audience growth (No. arXiv:2403.12838). arXiv. https://doi.org/10.48550/arXiv.2403.12838

Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial networks (No. arXiv:1406.2661). arXiv. https://doi.org/10.48550/arXiv.1406.2661

Gupta, A., Razavi, A., Toor, A., Gupta, A., Erhan, D., Shaw, E., Lau, E., Belletti, F., Barth-Maron, G., Shaw, G., Erdogan, H., Sidahmed, H., Nandwani, H., Moraldo, H., Kim, H., Blok, I., Donahue, J., Lezama, J., Mathewson, K., … Chen, Y. (2024). Veo 2. https://deepmind.google/technologies/veo/veo-2/

Holzinger, A., Saranti, A., Angerschmid, A., Finzel, B., Schmid, U., & Mueller, H. (2023). Toward human-level concept learning: Pattern benchmarking for AI algorithms. Patterns, 4(8), 100788. https://doi.org/10.1016/j.patter.2023.100788

Hunter, J. D. (2007). Matplotlib: A 2D graphics environment. Computing in Science & Engineering, 9(3), 90–95. https://doi.org/10.1109/MCSE.2007.55

Ideogram 3.0. (2025, March 26). https://about.ideogram.ai/3.0

Kampf, K., & Brichtova, N. (2025). Experiment with Gemini 2.0 Flash native image generation. Google Developers Blog. https://developers.googleblog.com/en/experiment-with-gemini-20- flash-native-image-generation

Karras, T., Laine, S., & Aila, T. (2019). A style-based generator architecture for generative adversarial networks (No. arXiv:1812.04948). arXiv. https://doi.org/10.48550/arXiv.1812.04948

Kätsyri, J., Förger, K., Mäkäräinen, M., & Takala, T. (2015). A review of empirical evidence on different uncanny valley hypotheses: Support for perceptual mismatch as one road to the valley of eeriness. Frontiers in Psychology, 6. https://doi.org/10.3389/fpsyg.2015.00390

Levine, T. R. (2020). Duped: Truth-default theory and the social science of lying and deception. University of Alabama Press.

Liu, Y., Zhang, K., Li, Y., Yan, Z., Gao, C., Chen, R., Yuan, Z., Huang, Y., Sun, H., Gao, J., He, L., & Sun, L. (2024). Sora: A review on background, technology, limitations, and opportunities of large vision models (No. arXiv:2402.17177). arXiv. https://doi.org/10.48550/arXiv.2402.17177

Lu, Z., Huang, D., Bai, L., Qu, J., Wu, C., Liu, X., & Ouyang, W. (2023). Seeing is not always believing: Benchmarking human and model perception of ai-generated images (No. arXiv:2304.13023). arXiv. https://doi.org/10.48550/arXiv.2304.13023 

Lualeperez. (2025). Lualeperez/coursera-introduction-to-deep-learning-with-keras: Coursera: Introduction to deep learning & neural networks with keras. https://github.com/lualeperez/coursera-introduction-to-deep-learning-with-keras

Meta. (2025, April 12). Misinformation Community Standards. Transparency Center. https://transparency.meta.com/policies/community-standards/misinformation/

Miller, E. J., Steward, B. A., Witkower, Z., Sutherland, C. A. M., Krumhuber, E. G., & Dawel, A. (2023). AI hyperrealism: Why ai faces are perceived as more real than human ones. Psychological Science, 34(12), 1390–1403. https://doi.org/10.1177/09567976231207095

OpenAI. (2025). Introducing 4o image generation. OpenAI. https://openai.com/index/introducing- 4o-image-generation

Raghavan, P. (2024, February 23). Gemini image generation got it wrong. We’ll do better. The Keyword; Google. https://blog.google/products/gemini/gemini-image-generation-issue/

Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv. https://doi.org/10.48550/ARXIV.2204.06125

Recraft AI. (2024, October 30). Recraft introduces a revolutionary AI model that thinks in design language. https://www.recraft.ai/blog/recraft-introduces-a-revolutionary-ai-model-that-thinks- in-design-language

Reve. (2025, March 24). Halfmoon is Reve Image—And it’s the best image model in the world. https://x.com/reveimage/status/1904211082870456824

​​Robins-Early, N. (2024, August 26). How did Donald Trump end up posting Taylor Swift deepfakes? The Guardian. https://www.theguardian.com/technology/article/2024/aug/24/trump-taylor- swift-deepfakes-ai

Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back- propagating errors. Nature, 323(6088), 533–536. https://doi.org/10.1038/323533a0

Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., & Ganguli, S. (2015). Deep unsupervised learning using nonequilibrium thermodynamics (No. arXiv:1503.03585). arXiv. https://doi.org/10.48550/arXiv.1503.03585

Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(56), 1929–1958. http://jmlr.org/papers/v15/srivastava14a.html

TikTok. (2024, April 17). Integrity and authenticity. TikTok Community Guidelines. https://www.tiktok.com/community-guidelines/en/integrity-authenticity

Top websites ranking—Most visited websites in April 2025 | Similarweb. (2025, April 1). [Data Analytics]. Similarweb; Similarweb. https://www.similarweb.com/top-websites/

Valyaeva, A. (2023, August 15). AI image statistics for 2024: How much content was created by Ai. Everypixel Journal. https://journal.everypixel.com/ai-image-statistics

Vukojičić, M., & Veinović, M. (2023, May). Imitation drawing: Can we spot the difference between AI and human generated drawing? https://doi.org/10.15308/Sinteza-2023-23-27

Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.-W., Chen, D., Yu, F., Zhao, H., Yang, J., Zeng, J., Wang, J., Zhang, J., Zhou, J., Wang, J., Chen, J., Zhu, K., Zhao, K., Yan, K., Huang, L., … Liu, Z. (2025). Wan: Open and advanced large-scale video generative models (No. arXiv:2503.20314). arXiv. https://doi.org/10.48550/arXiv.2503.20314

X. (2025, January). Authenticity. X Rules and Policies. https://help.x.com/en/rules-and- policies/authenticity

xAI. (2025, April 6). Grok image generation release | xAI. https://x.ai/news/grok-image-generation- release

Yang, C., Xiao, F., Cheng, F., Chen, H., Guo, H., Wei, M., Zhu, P., Zhao, Q., Lin, S., Zhao, Y., Lin, Z., Qing, Z., Kong, F., Zuo, F., Yan, J., Gui, L., Qi, L., Wang, S., Bi, S., … Yang, J. (2025). Seaweed-7b: Cost-effective training of video generation foundation model [Technical Report]. ByteDance. https://seaweed.video/

Appendix A

AI-Generated Image Dataset

The full dataset of images used in this research is publicly available to download on GitHub and the full set of AI-generated images used in the research is included below.


A diverse range of image types are included in the study, preventing  bias towards any one category, which could potentially skew the results. Differentiated image types also assess the ability to detect AI-generated images across different visual domains.

Table A1

Visual Categories used for Image Dataset

TypeNumber of
AI images
in dataset
Number of real
images in dataset
Nature/Landscape55
Multiperson55
Single person55
Celebrity/political figure55
Architectural/urban landscape55
Object55


For use in the evaluation, all images were scaled to a consistent resolution of 1024 × 768 pixels (approximately 0.8 megapixels) in order to control resolution as a potential confounding variable. This choice is also consistent with how most people interact with images online or through social media, where compression and lower resolution images are common.

Figure A1

Model: FLUX-pro (Black Forest labs)

Prompt: “Create a grainy TV quality photograph of Donald Trump talking at a political rally, supporters with trump 2024 signs behind him”

Figure A2

Model: FLUX-pro (Black Forest labs)

Prompt: “A photo of volunteers planting trees in Detroit, wearing green t-shirts.”

Figure A3

Model: Imagen 3 (Google)

Prompt: “A busy dive bar full of people with ambient lighting.”

Figure A4

Model: FLUX-pro (Black Forest labs)

Prompt: “A towering cliffside overlooking a busy highway ocean. A black and white lighthouse stands on the cliff. The sky with muted colors, city lights in the distance”

Figure A5

Model: Imagen 3 (Google)

Prompt: “New York street photograph taken from the road looking up, wide angle lens.”

Figure A6

Model: Imagen 3 (Google)

Prompt: “A closeup photograph of a baton handoff with elite runners, motion blur.”

Figure A7

Model: Recraft V3  (Recraft AI)

Prompt: “Vintage-style photograph of a worker in harsh weather, carrying a basket of fish over their shoulder, standing under heavy rain.”

Figure A8

Model: FLUX-pro (Black Forest labs)

Prompt: “A handwritten quote on a coffee shop blackboard that says ‘Life is a canvas, and you are the brush—paint your own masterpiece.’ Anna Sterling”

Figure A9

Model: FLUX-pro (Black Forest labs)

Prompt: “Nicole Kidman wearing a white dress in front of a media wall, holding an Oscar”

Figure A10

Model: Imagen 3 (Google)

Prompt: “Guys playing soccer in their backyard during the afternoon. Rough, patchy grass with dirt and weeds.”

Figure A11

Model: Imagen 3 (Google).

Prompt: “A woman hanging up a black and white photo in a messy art studio, many other photos hung up on the walls”

Figure A12

Model: Recraft V3 (Recraft AI)

Prompt: “Nighttime shanghai skyline with illuminated skyscrapers reflected in a river.”

Figure A13

Model: Imagen 3 (Google)

Prompt: “A bright Italian coastal town, drone shot”

Figure A14

Model: Recraft V3 (Recraft AI)

Prompt: “A bright mountain landscape reflected in a serene lake. White fluffy clouds in the sky.”

Figure A15

Model: FLUX-1.1-pro (Black Forest labs)

Prompt: “A dense forest with sunlight filtering through the trees.”

Figure A16

Model: FLUX-1.1-pro (Black Forest labs)

Prompt:  “A grey, soft focus photo of pale red poppies in a field.”

Figure A17

Model: FLUX-pro (Black Forest labs)

Prompt:  “Tom Brady, uniform 12, being lifted up by teammates at a football game.”

Figure A18

Model: Imagen 3 (Google)

Prompt: “A group of friends hiking through a forest trail on a snowy day. Wide angle deep depth of field”

Figure A19

Model: Imagen 3 (Google)

Prompt: “A man walking down the street bird poop on his face, disgusted look”

Figure A20

Model: Imagen 3 (Google) 

Prompt: “Create a photo of a lighthouse in Maine”

Figure A21

Model: Imagen 3 (Google)

Prompt: “A close-up of a vintage pocket watch on a rustic wooden table.”

Figure A22

Model: Recraft V3 (Recraft AI)

Prompt: “A pineapple sitting on a rough bench in a tropical setting.”

Figure A23

Model: FLUX-1.1-pro (Black Forest labs)

Prompt: “A long exposure photograph of rocks along a beach with gentle waves.”

Figure A24

Model: FLUX-1.1-pro (Black Forest labs)

Prompt: “A bee collecting pollen from a vibrant flower, its fuzzy body dusted with golden grains.”

Figure A25

Model: FLUX-pro (Black Forest labs)

Prompt: “Elton John performing live on stage with vibrant lighting.”

Figure A26

Model: FLUX-pro (Black Forest labs)

Prompt: “President Joe Biden giving a speech at the United Nations.”

Figure A27

Model: Imagen 3 (Google)

Prompt: “A person jogging along a beach side trail with palm trees during sunrise.”

Figure A28

Model: FLUX-1.1-pro (Black Forest labs)

Prompt: “A Buddhist monk wearing a brown robe, meditating peacefully in a moderately shabby temple with cracked painted concrete, zoomed out to show more of the temple.”

Figure A29

Model: FLUX-pro (Black Forest labs)

Prompt: “A steaming cup of coffee next to an open notebook with a pen.”

Figure A30

Model: FLUX-pro (Black Forest labs)

Prompt: “Artistic photograph of a bicycle parked beside a graffiti-covered wall.”

Appendix B:

Real Image Dataset

The full dataset of images used in this research is publicly available to download on GitHub and the full set of real photographs used in the research is included below. 

Images were selected based on the same categories described in Appendix A, and processed in a consistent manner. For real images, only photographs with robust EXIF data and a publication date prior to 2021 (the year advanced image generators became mainstream) were selected to use, as this reduces the probability of an AI-generated image being included in the dataset of real photographs. 

Figure B1

Details: Photograph of Kendrick Lamar, taken on February 7, 2013. This image was originally published on Wikimedia Commons by Merlijn Hoek. The image is licensed under Creative Commons BY-SA/GFDL.

Figure B2

Details: Photograph of a man standing in front of a bowl and looking towards the left, taken by Clem Onojeghuo on December 3, 2016. This image is free to use and available on Pexels.

Figure B3

Details: Photograph of a group of people enjoying a music concert, taken by Leah Newhouse. The image was uploaded on February 19, 2017, and is free to use. It is available on Pexels.

Figure B4

Details: Photograph of the Statue of Liberty during nighttime, taken by Pierre Blaché on March 22, 2019. This image is free to use and available on Pexels.

Figure B5

Details: Photograph of a cityscape view, taken by Aleksandar Pasaric on February 17, 2016. This image is free to use and available on Pexels.

Figure B6

Details: Photograph of a paintbrush in shallow focus, taken by Daian Gan on April 28, 2010. This image is free to use and available on Pexels.

Figure B7

Details: Close-up photograph of a pomegranate fruit, taken by Roman Odintsov on July 7, 2020. This image is free to use and available on Pexels.

Figure B8

Details: Photograph of a desert during nighttime, taken by Walid Ahmad on January 19, 2018. This image is free to use and available on Pexels.

Figure B9

Details: Photograph of Meryl Streep at the Opening Ceremony of the Tokyo International Film Festival. Taken on October 25, 2016 by Dick Johnson. This image was originally published on Wikimedia Commons. The image is licensed under Creative Commons BY-SA.

Figure B10

Details: Photograph of three persons sitting on stairs talking with each other, taken by Buro Millennial in Leiden, ZH, Netherlands. This image was uploaded on September 21, 2018, and is free to use. It is available on Pexels.

Figure B11

Details: Photograph of a group of women lying on yoga mats under a blue sky, taken by Amin Sujan on June 14, 2018. This image is free to use and available on Pexels.

Figure B12

Details: Photograph of a woman leaning back on a tree trunk using a black DSLR camera during the day, taken by David Bartus on October 7, 2017. This image is free to use and available on Pexels.

Figure B13

Details: Photograph of Saint Basil's Cathedral on New Year's Eve in Moscow, taken by Y Nakanishi on December 31, 2014. The photo was uploaded to Flickr and is licensed with some rights reserved.

Figure B14

Details: Photograph of a white soccer ball, taken by Aphiwat Chuangchoem. This image was uploaded on March 28, 2017, and is free to use. It is available on Pexels.

Figure B15

Details: Photograph of waterfalls in the middle of green trees, taken by Greg Galas on June 12, 2019. This image is free to use and available on Pexels.

Figure B16

Details: Photograph of selective-focus red fruits with snow taken by Nadine Wuchenauer on December 26, 2018. This image is free to use and available on Pexels.

Figure B17

Details: Photograph of President Barack Obama hosting a press conference at the Pentagon in Washington, D.C., on August 4, 2016. Taken by Air Force Tech. Sgt. Brigitte N. Brantley. This image is in the public domain and was uploaded to Flickr.

Figure B18

Details: Photograph of a crowd dancing in a blue-painted enclosure, taken by Maurício Mascaro on March 16, 2018. This image is free to use and available on Pexels. Licensed under Creative Commons.

Figure B19

Details: Photograph of a man in a white T-shirt sitting on a brown rock formation, taken by Tima Miroshnichenko on November 10, 2020. This image is free to use and available on Pexels.

Figure B20

Details: Photograph of a woman wearing a black sleeveless dress holding white headphones during daytime, taken by Tirachard Kumtanom on June 4, 2017. This image is free to use and available on Pexels.

Figure B21

Details: Photograph of a gray spiral building taken from a low angle, taken by Li Jin on October 31, 2019. This image is free to use and available on Pexels.

Figure B22

Details: Close-up photograph of a ukulele, taken on May 25, 2012. This image is free to use under CC0 and is available on Pexels.

Figure B23

Details: Black wooden fence on snow field at a distance of black bare trees taken on February 16, 2012. This image is free to use under the CC0 license and is available on Pexels.

Figure B24

Details: Photograph of a mountain under a cloudy sky, taken by Evgeny Tchebotarev on January 1, 2014. The image is free to use and is available on Pexels.

Figure B25

Details: Photograph of Simone Biles at the 2016 Olympics all-around gold medal podium in Rio de Janeiro. Taken on August 9, 2016, at 18:46 by Agência Brasil Fotografia. This image was originally published on Wikimedia Commons. and is licensed under Creative Commons BY.

Figure B26

Details: Photograph of President Ronald Reagan speaking at a rally for Senator Durenberger on February 8, 1982. Taken by Michael Evans and part of the Ronald Reagan Library collection (C6289-25). This image is in the public domain, sourced from the National Archives.

Figure B27

Details: Photograph of a person standing in front of a brown crate, taken by Clem Onojeghuo on December 3, 2016. This image is free to use and available on Pexels.

Figure B28

Details: Photograph of a man in a black suit riding a bicycle down the street, taken by Andrea Piacquadio. This image was uploaded on January 28, 2019, and is free to use. It is available on Pexels.

Figure B29

Details: Aerial photograph of city buildings near Honolulu, taken by Jess Loiterton on May 21, 2020. This image is free to use and available on Pexels.

Figure B30

Details: Photograph of a pink and white keychain, taken by Marcin Szmigiel on May 23, 2015. This image is free to use and available on Pexels.

Appendix C

F-table of Critical Values

Df1 represents degrees of freedom between groups and Df2 represents degrees of freedom within groups. An F statistic greater than its critical value results in the rejection of the null hypothesis (there is no difference between groups).

Table C1
F-table of Critical Values for α = 0.05

DF1/DF21234567891011152024304060120
2161199215224230233236238240241243245248249250251252253254
318.51919.119.219.319.319.319.319.319.419.419.419.419.419.419.419.419.419.5
410.19.559.289.129.018.948.898.858.818.798.748.78.668.648.628.598.578.558.53
57.716.946.596.396.266.166.096.0465.965.915.865.85.775.755.725.695.665.63
66.615.795.415.195.054.954.884.824.774.744.684.624.564.534.54.464.434.44.37
75.995.144.764.534.394.284.214.154.14.0643.943.873.843.813.773.743.73.67
85.594.744.354.123.973.873.793.733.683.643.573.513.443.413.383.343.33.273.23
95.324.464.073.843.693.583.53.443.393.353.283.223.153.123.083.043.012.972.93
105.124.263.863.633.483.373.293.233.183.143.073.012.942.92.862.832.792.752.71
154.63.743.343.112.962.852.762.72.652.62.532.462.392.352.312.272.222.182.13
204.383.523.132.92.742.632.542.482.422.382.312.232.162.112.072.031.981.931.88
304.183.332.932.72.552.432.352.282.222.182.12.031.941.91.851.811.751.71.64
404.083.232.842.612.452.342.252.182.122.0821.921.841.791.741.691.641.581.51
6043.152.762.532.372.252.172.12.041.991.921.841.751.71.651.591.531.471.39
1203.923.072.682.452.292.182.092.021.961.911.831.751.661.611.551.51.431.351.25
3.8432.62.372.212.12.011.941.881.831.751.671.571.521.461.391.321.221






Appendix D

MLP Training Script

The following Python script implements a multilayer perceptron for predicting participant accuracy based on demographic and confidence variables. This script was based on an original implementation as a part of the Introduction to Deep Learning and Neural Networks with Keras,  a course offered by IBM (lualeperez, 2025). It utilizes the TensorFlow and Keras frameworks (Abadi et al., 2016) and modeled after the original neural network design (Rumelhart et al., 1986). The feature importance analysis is adapted from Breiman's variable importance measures (Breiman, 2001). 


The script employs standard practices (Srivastava et al., 2014), uses Matplotlib (Hunter, 2007).

# Import statements set up necessary tools and libraries. These tools allow the implementation of the model as described above.
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.metrics import mean_squared_error, r2_score
import tensorflow as tf
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Dropout
from tensorflow.keras.callbacks import EarlyStopping

# A random seed is used to set up both numpy and tensorflow. Because the seed is fixed, this ensures reproducibility by future researchers, and similar results should be produced each time the script is run.
np.random.seed(42)
tf.random.set_seed(42)

# To analyze a dataset, data is imported from a file set as “demographics.csv” which has the nine input demographic factors as the first nine columns. Factors were converted to quantitative values before analysis.
print("Loading data from demographics.csv...")
data = pd.read_csv('demographics.csv')
print(f"Dataset shape: {data.shape}")

# Ensures that there were not any mistakes in the data formatting. (see previous comment)
print("\nData types of each column:")
print(data.dtypes)

# Check again for mistakes
print("\nMissing values in each column:")
print(data.isnull().sum())

# the rest of the script
X = data.iloc[:, :9]  # First 9 columns are the demo. Factors for each participant.
y = data.iloc[:, 9]   # The 10th column is each participant’s accuracy on the evaluation.

# Splits data into training and validation sets. A validation set ensures the model is actually generalizable in the real world. A fairly large 20% validation set is used in this case. 
X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)

# Standardize the input features using z scores!
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_val_scaled = scaler.transform(X_val)

# Create model is created here
print("Creating the model...")
model = Sequential([
    Dense(16, activation='relu', input_shape=(9,)),  # This is the first hidden layer, which has 16 neurons and uses ReLU activation.
    Dropout(0.2),  # Randomly turns off some neurons, which prevents overfitting, a concern as we only have 280 participants as training examples.
    Dense(8, activation='relu'),                     # Hidden layer 2, with 8 neurons.
    Dense(1, activation='sigmoid')                   # Output neuron, which compresses the values into a probability between 1 and 0 using a sigmoid function. Because probabilities are values on a scale from 0 to 1, the sigmoid simply normalizes the model’s prediction into an easier to understand metric.
])

# This model uses an adam optimizer for training, which is standard practice. Instead of optimizing a loss function, this model is designed to minimize the difference between predictions and actual participant’s accuracy.
model.compile(
    optimizer='adam',
    loss='mean_squared_error',  # MSE for regression
    metrics=['mae']
)

# Print function
model.summary()

# This sets up early stopping to avoid overfitting. The script will wait 30 training cycles without improvement and keep track of the validation loss.
early_stopping = EarlyStopping(
    monitor='val_loss',
    patience=30,
    restore_best_weights=True
)

print("\nTraining the model...")
history = model.fit(
    X_train_scaled, y_train,
    epochs=200,
    batch_size=16,  # Small batch size as this is a small dataset. If this research is reproduced with an extremely large sample size increasing this from 16 might help speed up training.
    validation_data=(X_val_scaled, y_val),
    callbacks=[early_stopping],
    verbose=1
)

# Evaluate the model’s performance using the validation dataset.
print("\nEvaluating model performance...")
y_pred_val = model.predict(X_val_scaled).flatten()
val_mse = mean_squared_error(y_val, y_pred_val)
val_rmse = np.sqrt(val_mse)
val_r2 = r2_score(y_val, y_pred_val)

print(f"Validation MSE: {val_mse:.4f}")
print(f"Validation RMSE: {val_rmse:.4f}")
print(f"Validation R²: {val_r2:.4f}")

# Make predictions using the entire dataset, in order to compare to actual values.
X_all_scaled = scaler.transform(X)
y_pred = model.predict(X_all_scaled).flatten()

# Create a plot of the training history to observe how well the model learns.
plt.figure(figsize=(12, 5))
plt.subplot(1, 2, 1)
plt.plot(history.history['loss'], label='Training Loss')
plt.plot(history.history['val_loss'], label='Validation Loss')
plt.title('Model Loss During Training')
plt.xlabel('Epoch')
plt.ylabel('Loss (MSE)')
plt.legend()
plt.grid(True)

plt.subplot(1, 2, 2)
plt.plot(history.history['mae'], label='Training MAE')
plt.plot(history.history['val_mae'], label='Validation MAE')
plt.title('Model MAE During Training')
plt.xlabel('Epoch')
plt.ylabel('Mean Absolute Error')
plt.legend()
plt.grid(True)

plt.tight_layout()
plt.savefig('train.png')
plt.show()

# Create a regression analysis of actual values vs values predicted by the model.
plt.figure(figsize=(10, 6))
plt.scatter(y, y_pred, alpha=0.5)
plt.xlabel('Actual Values')
plt.ylabel('Predicted Values')
plt.title('MLP Prediction Performance')
plt.grid(True)
plt.savefig('regression.png')
plt.show()

print("\nAnalysis completed!")






Appendix E

Survey Materials

Data was collected using Google Forms. The survey and questions are publicly available on Forms and the questions are included below. Participants were presented with the following information before consenting.

Introduction

I am a student at Morse High School, and I am in a class that requires me to complete a research study. My project is to investigate how well individuals can distinguish between AI-generated images and real photographs.

You have been selected as a possible participant because you are part of community outreach efforts. There are no specific exclusionary criteria; anyone aged 18 or older is welcome to participate.

Please read this form before agreeing to participate in this study.

Purpose of the Study

The purpose of the study is to examine how effectively people can differentiate between AI-generated images and real photographs, and to understand whether receiving feedback improves their ability and confidence in subsequent attempts.

I intend to present the findings of this study in a written academic paper submitted to the College Board, aligning with their terms. I may submit the results for publication in a student research journal.

If you participate in this study, you will be asked to complete an online survey consisting of two sets of images. In the first set, you will view 30 images and classify each one as either real or AI-generated. After completing the first set, you will receive feedback on your responses. Then, you will evaluate a second set of 30 images in the same manner. Throughout the study, you will also be asked to rate your confidence levels using three scale questions ranging from 1 to 6. The entire study will take approximately 5–10 minutes to complete.

Risks and Benefits

The study involves minimal risks. There are no foreseeable psychological or physical risks, but as with any research, there may be unknown risks that are currently unforeseeable.

There are no direct benefits, gifts, or payments for participating in this study. Your participation will contribute to research that may help improve understanding of how people perceive AI-generated content.

Confidentiality

This study is confidential, and I will not be retaining or reporting any information about your identity. Findings from the study will not include any information that would make it possible to identify you, and results will only be reported in aggregate.

The records of this study will be kept strictly secure. Data will be stored on a password-protected computer accessible only by me. No audio or video recordings will be made.

By agreeing to participate, you consent to the inclusion of the study findings in any reports or publications.

Questions and Concerns

You have the right to ask questions about this research study and to have those questions answered before or after your participation. If you have any further questions about the study, feel free to contact me, Declan Wright, at declan.wright@rsu1.org.

A summary of the results of the study can be sent to you upon request.

Survey Questions

Email
(Note: Email collected automatically for potential result summary distribution, not linked to responses in analysis)

Certification

I have read the above terms and consent to voluntary participation in this research.

Background Information

What year were you born? (e.g., 1986)

What is your gender?

Male

Female

Other:

What is your current education level?
(Dropdown menu with options ranging from "No formal education" to "Doctorate (PhD or equivalent)" and "Vocational or technical training")

Have you ever encountered an image online that you suspected was AI-generated?

Yes

No

In which state do you currently reside? If you are outside the United States, please select "International."
(Dropdown menu with options for all US states and "International")

Approximately how many hours per day do you spend on your smartphone? (Average daily screen time for the last full week, rounded to the nearest hour)
(Dropdown menu with options from 0 hours to 20 hours)

Are there any other factors that might be significant to your ability to discern AI-generated and real photographs? (e.g., photographer, artist, familiarity with AI tools, tech industry employee) Please share them now.

Before answering any questions, how confident are you in your ability to tell the difference between real and AI-generated images?
(Scale: 1 = Not confident at all, 6 = Extremely confident)

Image Classification Set (Round 1)

(Participants were shown 30 images, one at a time)

11-40. For each image presented (Image 1 through Image 30):

Do you think this image is:

AI-generated

Real photograph

Feedback Section

(Participants were shown the following instructions and the correct answers for images 11-40)

Answers to the first set of images: Review your mistakes and analyze image features as necessary. DO NOT change answers to the previous section under any circumstances.

(Correct classifications for the 30 images were displayed here)

After reviewing the correct answers from the first set of images, please rate your confidence in distinguishing between AI-generated images and authentic photographs.
(Scale: 1 = Not confident at all, 6 = Extremely confident)

Image Classification Set (Round 2)

(Participants were shown the second set of 30 images, one at a time)

42-71. For each image presented (Image 31 through Image 60):

Do you think this image is: 

AI-generated

Real photograph

Closing

Please rate your confidence in distinguishing between AI-generated images and authentic photographs after completing the second set of images.
(Scale: 1 = Not confident at all, 6 = Extremely confident)

How helpful was the feedback you received after completing the first set of images in improving your ability to distinguish between AI-generated images and authentic photographs?
(Scale: 1 = Not helpful at all, 6 = Extremely helpful)

Type the word "blue" before continuing.

(End of Survey)






Appendix F

Small-Scale Evaluation with Newer Image Models

Following the completion of the initial round of data collection for this study, several newer and more advanced AI image generation models were released by developers. These models represent strong improvements in photorealism and detail compared to those used in the main evaluation dataset. To assess whether the main finding, that human detection accuracy had fallen to chance levels, remains true against these state-of-the-art models, a small-scale follow-up evaluation was conducted in early 2025.

Table F1
Models Used in Image Generation

ModelDeveloperRelease DateImages Generated
Imagen3-002Google DeepMindFebruary 5, 202540%
Recraft V3*Recraft AIOctober 30, 202437%
FLUX 1.1 [pro] RawBlack Forest LabsNovember 6, 202413%
AuroraxAI Corp.December 9, 202410%

*Includes Photoreal, Hard Flash, and Black & White finetuned versions of the model.

The evaluation format mirrored the main study, presenting participants with a total of 60 real photographs and synthetic images from the newer models above. For simplicity, images were not split into multiple feedback groups, and reduced demographic information was collected. A small group of participants (n = 15) completed this follow-up.

The results from this small-scale evaluation indicated that participant accuracy remained near chance levels (mean accuracy of 48.5%). This finding aligns closely with the original observations of the main study (48.7% and 48.3% accuracy). Sample t-tests result in a p value of 0.90. While a small sample size does add some uncertainty to this finding, the effect size from the sample also shows us something interesting: if a difference were to truly exist in the broader population, we would require a sample size of around 1.6 million people to confidently observe that difference, which highlights how minuscule even a true change would be. In other words, the critical threshold of human ability to detect AI-generated images at random levels persists, and is potentially solidified with the very latest advancements in diffusion model technology available at the time of the follow-up. It is important to note that the field continues to evolve at an extremely fast pace, with even newer models and architectures (such as image generation powered by large language models) being released regularly even after this follow-up was conducted.

Interestingly, while the mean accuracy remained stable, there is an observed change in the variance of participant scores occurred between the main study and this follow-up. The overall standard deviation (of all 60 images together) decreased from 7.5% in the original research to 5.4% in the follow-u. To formally evaluate the difference in variances, Levene’s Test was used. For each observation in each group, Levene's test calculates the deviation from its group's median, then performs an ANOVA on these values. This resulted in a W statistic of 2.58.

As shown in Figure F1, this observed W statistic falls below the critical values required for statistical significance at commonly used alpha levels (α = 0.05 and α = 0.10). Therefore, we do not have enough evidence to reject the null hypothesis that the variances are equal; the observed difference is not statistically significant.

Figure F1
Levene’s Test F Distribution With Selected Values

Figure or table reproduced from the original research paper


It is notable that the observed W statistic of 2.58 approaches the critical value for α = 0.10. While failing to meet the threshold, this proximity might suggest a potential trend towards reduced variability in detection accuracy when participants face the very latest models. A lower variance might suggest that images created by the next generation of image models are more consistently challenging, due to the tighter association around the chance-level mean. This hypothesis is speculative, but may be worth additional research, particularly given the limitation of the very small sample size in this follow-up study. Such a small sample is highly susceptible to random chance, and a larger sample would be required to evaluate this possibility.

Overall, this follow-up builds further evidence to support the main study: human accuracy in detecting state-of-the-art AI-generated images is currently indistinguishable from random chance. The observed decrease in score variance and its proximity to significance thresholds deserves mention, but does not suggest any new understanding within itself, other than that the field's rapid advancement requires continued evaluation of human limits against AI advancement.