A study led by researchers associated with Stanford University has raised concerns about the reliability of artificial intelligence detection tools used to identify whether students have used ChatGPT or other generative AI systems in their academic writing.
The study examined essays written by non native English speakers preparing for the Test of English as a Foreign Language, commonly known as TOEFL. Researchers evaluated 91 essays using seven widely used AI detection systems. The results showed that the detectors classified 61.22 percent of the essays as AI generated on average. At least one of the seven detectors flagged 89 of the 91 essays, equivalent to about 97 percent of the sample.
The findings have important implications for schools, colleges and universities that use automated tools to assess whether students have relied on generative AI. A positive result from an AI detector does not necessarily establish that a student used ChatGPT or another AI system. Instead, the study suggests that certain characteristics of human writing, particularly writing by people who use English as a second language, can resemble patterns associated with AI generated text.
The research was conducted by Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou. The paper, titled GPT Detectors Are Biased Against Non Native English Writers, examined the performance of AI detection systems on writing samples from both native and non native English speakers.
One of the central issues identified by the researchers is the way many AI detectors analyse language. Such systems can use statistical characteristics of writing, including the predictability of words and sentence structures. One commonly discussed measure is perplexity, which is used to estimate how predictable a sequence of words is to a language model.
Human writing that uses a wide range of vocabulary and more varied sentence structures may appear less predictable to such systems. By contrast, writing that relies on familiar words, straightforward sentence patterns and commonly used grammatical structures can appear more predictable.
This can create difficulties for students who are learning or using English as a second language. Such students may naturally use vocabulary and sentence structures that they are more confident with. Their writing may therefore contain fewer unusual words or complex constructions, even when the work is entirely their own.
The Stanford researchers found a significant difference between the results for the TOEFL essays and essays written by US born eighth grade students. The AI detectors performed much more accurately on the latter group, while the non native English writing samples were frequently classified as AI generated.
The researchers said the results raise questions about whether AI detection systems can provide an objective assessment of authorship. If an automated detector incorrectly identifies human writing as AI generated, students could potentially face additional questioning or disciplinary action despite not having used generative AI.
This concern is particularly relevant as educational institutions increasingly look for ways to address the use of ChatGPT and other generative AI tools in assignments. Since the widespread availability of generative AI, schools and universities have been exploring software designed to identify potentially AI generated content.
However, the Stanford research indicates that detection scores have limitations. A detector does not observe the writing process itself. Instead, it analyses the final text and searches for characteristics that its model associates with machine generated language.
This distinction is important when interpreting an AI detection result. A high AI probability score should not automatically be treated as evidence that a student used ChatGPT. The research suggests that the characteristics of a student's language background can influence the outcome.
The study also examined whether the detectors could be bypassed. Researchers reported that simple prompting techniques could modify AI generated text in ways that reduced the effectiveness of some detection systems. This creates another challenge because a tool that can incorrectly flag human writing while also being circumvented by modified AI generated writing may not provide a reliable standalone method for determining authorship.
The researchers therefore cautioned against relying heavily on AI detectors in educational settings. Stanford's Human Centered AI report on the study said the researchers recommended particular caution when these tools are used in environments with large numbers of non native English speakers.
The issue goes beyond language proficiency. Academic institutions must also consider how evidence is collected when investigating suspected AI use. A student's writing history, drafts, notes, document revision history and discussions with teachers can provide additional context about how an assignment was produced. An automated detector score represents only an analysis of the submitted text.
The study also highlights a broader challenge facing developers of AI detection technology. As generative AI systems become more sophisticated, distinguishing between human and machine generated text using only statistical patterns becomes increasingly difficult. Detection systems must therefore be evaluated across different languages, educational backgrounds and writing abilities rather than relying on a narrow set of test samples.
For students, the findings underline the importance of retaining evidence of their writing process. Drafts, research notes, outlines and earlier versions of assignments can help demonstrate how a piece of work developed over time if questions about authorship arise. Such records do not automatically prove authorship, but they can provide useful context during an academic review.
The findings also have implications for non native English speakers who may already face challenges in academic environments. If an automated system disproportionately flags writing because it is relatively simple or predictable, students could face additional scrutiny based on their language style rather than the actual method they used to produce their work.
James Zou and his colleagues said the current generation of AI detectors needs further evaluation and improvement before it can be relied upon for high stakes decisions. The researchers suggested that developers should investigate more sophisticated approaches rather than depending primarily on measures such as perplexity. They also discussed other possible approaches, including watermarking systems that could embed signals into AI generated content.
The study does not mean that AI detection technology has no potential use. Rather, it highlights the limitations of treating automated detection as definitive evidence. Detector performance can vary depending on the type of writing, the language background of the author, the detector being used and the way the text has been produced or edited.
The research also provides an important distinction between detecting AI generated text and proving academic misconduct. Even when a detector identifies a passage as likely AI generated, that result alone does not establish when or how the text was produced. Human review and additional evidence may therefore be important when academic institutions investigate suspected violations.
The Stanford findings were published in 2023, shortly after the rapid adoption of generative AI tools such as ChatGPT. Since then, AI detection technology has continued to develop, meaning the specific performance of individual tools can change over time. The study nevertheless remains relevant to the broader question of whether automated systems can reliably determine authorship from writing style alone.
For educational institutions, the research highlights the need for careful interpretation of AI detector results, especially when decisions could affect a student's academic record. For students, it shows why an AI detection score should not automatically be viewed as proof that an assignment was generated by ChatGPT.
The central finding from the Stanford research is that human writing by non native English speakers can be incorrectly classified as AI generated. Among the 91 TOEFL essays examined, 61.22 percent were classified as AI generated on average by seven detectors, while 89 essays were flagged by at least one detector. These findings have prompted researchers to question the reliability and fairness of using AI detection scores as the sole basis for academic judgments.





