Exclusive: Is AI writing detection reliable? Testing three popular tools

In recent times, with the increasing calls for protecting and prioritizing human-authored content, the identification of AI-generated text has become a familiar challenge for publishers, educators, and other professionals.

The emergence of AI writing detection software in this context has risen to meet this challenge, with some tool developers claiming up to a 99% accuracy rate in distinguishing between human and AI writing.

However, many legal experts caution that conclusions from AI writing detection software must be handled with care. Relying on AI detection scores to determine if an article is AI-generated could lead to serious legal repercussions for educators, publishers, and employers.

Experts in AI — including those who have developed detection tools — acknowledge that these services are intended to provide useful clues but are not infallible. In fact, they seem to share a common blind spot.

Research on the effectiveness of AI detection tools has been ongoing, with studies and experts emphasizing a key point: human involvement is still needed in the evaluation process.

In a study published in 2026 by the International Journal for Educational Integrity (IJEI) based in Australia, researchers noted that while detection tools can offer useful initial warnings, they should not be the sole evidence for decision-making with significant consequences. Detection results should be integrated into a broader assessment strategy.

Researchers and industry experts also express concerns about the accuracy of these detection tools, fearing they might mistakenly flag human writing as AI-generated.

A spokesperson from GPTZero, an AI startup based in New York, told the Epoch Times, “Detection results are just a signal, not a conclusion. It should spark further investigation rather than concluding the investigation.”

In professional fields, two things are more critical than the detection tools themselves.

The spokesperson stated, “Firstly, an organization needs to establish a written AI use policy before considering detection tools. These tools can help enforce standards but cannot replace standards themselves.”

“Secondly, focus on patterns rather than isolated cases and integrate detection results with process evidence.”

Within a two-week period, the Epoch Times tested three leading AI writing detection tools – Pangram, GPTZero, and Originality.AI. The testing results were eye-opening.

Using the same writing samples ranging from 500 to 2000 words in length for each platform’s free version, the samples included three types: 100% AI-written, 100% human-written, and mixed-mode with AI and human writing ratios of 25:75, 50:50, and 75:25, respectively.

The 100% AI-written samples were created by large language models ChatGPT and Claude to determine if certain automated writing styles are more easily detectable than others.

In the Epoch Times test, all three detection services easily identified the 100% AI-generated text samples.

Pangram and GPTZero provided a 100% probability score for AI generation, while Originality.AI believed the text at least had a 40% chance of being AI-generated.

During testing, the three detection tools performed similarly in identifying 100% human-written samples. However, when mixed writing samples were introduced, the situation became more intriguing.

Pangram showed the most significant variations in detection results for human-machine mixed writing samples. An originally 100% human-written text, after being edited to include 25% AI-generated content, showed an 87% ratio of AI-generated text in the new text.

When the human writing and AI-generated ratio was reversed to 25:75, Pangram scored the test as 100% AI writing.

In the testing, GPTZero and Originality.AI also considered the 25:75 mixed writing sample ratio as a tipping point. These two programs showed the most reliable results in detecting texts with that ratio, regardless of it being the original or reversed ratio. However, they performed poorly in detecting texts with a 50:50 ratio. Pangram had the highest accuracy rate in detecting texts with a 50:50 ratio among all the mixed writing samples.

However, the issue lies in the use of AI editing software such as Grammarly, which could influence the detection scores.

American author Mia Ballard denied using AI to create the novel “Shy Girl” in 2025. Previously, reports suggested that due to suspicion of heavy AI usage, she parted ways with the French publisher Hachette. Ballard suggested her editorial staff might have used AI during the editing process, but the reports only mentioned suspicion without providing conclusive evidence of Ballard using AI to create the novel.

The controversy over solely using AI for editing articles being flagged as violations by detection services has been ongoing. However, Jonathan Gillham, CEO of Originality.AI, stated that this scenario is entirely plausible.

“If I’m using Grammarly to edit my article, that includes a lot of AI technology… If you use AI tools to assist in editing, it imprints that information into the editing process,” explained Gillham to the Epoch Times.

This raises concerns for those who believe that solely using AI tools for editing can pass scrutiny by AI detection services. However, Gillham pointed out that there are some nuances in this regard.

Gillham noted that merely correcting punctuation and spelling errors might not trigger the detection tools, but accepting modification suggestions could potentially be identified. He believed that when someone accepts suggestions to rephrase paragraphs and sentences, these changes leave similar traces in the document, which AI detection tools may recognize.

Gillham suggested that when someone writes books, articles, or papers and uses AI tools for editing and modification, the work might be labeled as AI-generated.

This is why Gillham believed that it is crucial for companies to establish an “AI tolerance policy.”

“For example, if I work with someone who uses AI in their writing, and the company has a zero-tolerance policy for AI, then I can’t use auxiliary tools like Grammarly,” he stated.

To validate this theory, a 1,500-word human-written article was directly inputted into these three detection tools in the first round of testing with no modifications. Then, after making edits based on Grammarly suggestions, the article was once again presented to these three detection tools to see if it would generate AI writing scores.

In the initial testing round, none of the samples were detected as AI writing. In the second round of testing, Originality.AI and Pangram did not detect any non-human influence, while GPTZero detected 1% and noted this article as AI-generated.

Companies or publishers must specify the extent to which they can allow the use of AI in writing, crucial for avoiding mismatched expectations. Clarity in this matter can prevent companies from assuming unnecessary responsibilities. Gillham emphasized that the crux of the issue lies in the lack of transparency.

“AI detection tools are part of the discussion. On a large dataset, they are very effective tools, helping us understand how much AI content an organization is creating,” he said.

When asked if he believed AI detection tools could serve as the final basis for decision-making, Gillham firmly stated that no tool’s accuracy rate can reach 100%. Moreover, when dealing with mixed content, questions arise about what constitutes human and AI elements.

Kyle Szives, founder and senior software engineer at ANTLR Interactive in Pennsylvania, also agreed with this perspective.

“AIs in writing detection tools are technically unable to definitively identify the author’s identity. From a macro perspective, writing detection tools typically classify a piece of text as AI-generated or non-AI-generated by looking for common statistical signals in machine-generated texts,” explained Szives.

Having developed AI-assistive products, Szives stated that the presence of AI “probability or signal” does not conclusively prove that someone has used the technology in their writing.

“If a piece of human-written text is structurally rigid, formulaic, or predictable, unusually short, shows heavy editing traces, or follows a clear yet consistent professional style, it might be tagged as AI-generated text,” he said.

Szives pointed out that the same can happen in reverse when extensively rewriting or merging text generated by AI with human-written text. In such cases, there could be misjudgments identifying it as “non-AI generated.”

Gillham stated that AI writing detection tools are trained based on human and AI writing patterns, but their functioning is not significantly different from systems like ChatGPT — both involve pattern recognition.

When asked how detection tools determine AI scores, Gillham pointed out, “The disconcerting answer is, the AI system is a black box.” It’s the danger herein.

“I think there’s some misunderstanding about AI detection from outside. Tests show high accuracy rates of these detection tools, but some interpret their accuracy as 100%,” he said.

Undoubtedly, there is a pressing demand in the market for AI detection services.

Analysis by researchers from Originality.AI scrutinizing books on Amazon revealed suspicions of AI-generated content in 82% of herbal remedy books, 77% of success-related self-help books, and 63% of religious books.

These findings were supported by other research. In July this year, researchers from Stony Brook University, Columbia Law School, and the University of Michigan released their study results. They examined 14,419 self-published genre novels sold on Amazon between 2023 and 2026. Using AI detection tools, they found that 20% of the novels contained “substantial AI text.” Non-AI text books accounted for only 37% of the test group.

“When there are doubts about an author’s identity, other evidence should be considered, such as drafts, revision logs, source annotations, metadata, and the authors themselves,” Szives stated.

Some American lawyers emphasize that relying solely on AI detection tools when making critical decisions regarding employees, clients, or contracts could pose potential risks.

“If a company primarily relies on detection tool scores to dismiss employees, terminate contracts with contractors, or refuse job deliverables, it might make significant decisions based on a potentially faulty tool,” said Michael McCready, founder of McCready Law.

McCready mentioned the action taken by the Federal Trade Commission (FTC) against the AI detection tool Workado last year. The FTC accused Workado of claiming to clients that its accuracy rate was as high as 98%, while an independent investigation cited in the case found the tool’s accuracy rate to be about 53% on general content.

“One major issue is discrimination. If AI writing detection tools disproportionately flag certain groups and employers rely on these detection results for recruitment, disciplinary actions, or dismissal decisions, the company could face legal risks,” McCready stated.

Companies cannot evade responsibility by stating “the software told us to do it.”

“The ultimate decision-making authority still rests with the company. If a supplier has made unverified claims about the accuracy or reliability of the tool, they may face additional liability,” he said.

Alan Heimlich, President of Heimlich Law, concurred with this view.

“One of the biggest risks is that companies could mistake probability for evidence. Because in any case, AI writing detection tools cannot conclusively prove an author’s identity, and firmly believing in the accuracy of these technologies may lead companies to shoulder unwarranted responsibilities,” he said.

Gillham also emphasized that when AI detection scores unexpectedly emerge, it’s vital not to act hastily.

“We believe detection ratings should serve as a starting point for discussion whenever possible,” he stated.

Overall, the landscape of AI writing detection tools presents complex challenges and nuances, requiring a balanced approach that incorporates human judgement alongside technological assistance.