Did AI write this? It’s getting harder to tell
When Max Spero reads something online, he trusts his intuition to tell him whether it was written by AI even before he runs it through his detection software.
The former Google software engineer spent years working in machine learning, so his instinct is more nuanced than most. It’s not the classic signifiers of large language models he looks for — the em dashes, the hallucinations and the smoothly pleasant tone. Instead, it’s a sense of emptiness. “I look for information density,” he says. “When a person writes, every word has intention behind it. When AI generates text, it does not.”
Over the past year, Spero has become something of a hero to the small but determined group of self-appointed linguistic detectives calling out what they believe is the deceitful use of AI.
When you see a headline claiming a third of new web pages were written using AI, or a fifth of Kindle ebooks, the researchers will often have used Spero’s AI detector, Pangram. It was at the centre of the Shy Girl scandal, in which claims of AI led to the novel being removed from bookstores, despite author Mia Ballard denying she used the technology. This summer, it was used to accuse Trinidadian Jamir Nazir of submitting an AI story to the prestigious Commonwealth Short Story Prize.
Texts published by judges, consultants, lawyers and newspapers (including The New York Times, the FT* and The Wall Street Journal) have been flagged by AI detectors. Many who use the detection software say they are not opposed to AI itself, but dishonesty around its use. AI content passed off as human “pollutes the commons”, says Chris Best, the chief executive of publishing platform Substack, which is working with Pangram to scan its newsletters.
Spero has occasionally waded in on social media, posting screenshots and offering to run text through his model.
He and co-founder Bradley Emi, a fellow Stanford graduate and former machine learning scientist at Tesla, created Pangram in 2023 because they were worried about what would happen if we could no longer distinguish machine content from our own. Their pitch is that in the age of AI, authenticity is more valuable than ever.
But in the midst of this cat-and-mouse game are those who are less convinced that we can draw a clean line between human and AI text. Universities including Oxford, Cambridge and Harvard, which all support the use of generative AI by students, do not endorse the use of commercial AI detectors to scan work.
Writers point out that the tics commonly attributed to AI all appear in human writing and that both detection software and gut feelings can be wrong. Many worry about the implications of a false accusation. One person who works in the tech sector says the battle for total certainty was lost the moment OpenAI’s ChatGPT launched in late 2022.
Some of the writers accused of using AI are pushing back too, either claiming there is nothing wrong with using AI, or saying that paranoia around synthetic content has led to witch-hunts. Nazir has repeatedly said that he did not use AI to write his story, The Serpent in the Grove, which won the final Commonwealth prize. He compares AI and Pangram to Eris, the goddess of discord who threw down an apple and sparked the Trojan War.
“We trained this thing on all of our language,” he tells the FT. “And now it is impossible to trust each other.”
Arguments over authorial provenance are nothing new. The Roman poet Martial wrote a verse to a poet he accused of kidnapping his words (plagiario) in the first century. And for years, authors have delegated their work to ghostwriters, research assistants and editors without acknowledging their input.
The difference now is volume. The speed and ease with which AI can generate text means it is reasonable to assume that a lot of what we see online may have been artificially generated.
AI detectors like Pangram, GPTZero, Winston AI or Copyleaks let users feed in suspicious extracts and produce a percentage score showing how much of the text they estimate was generated by AI. Limited free versions are available while subscriptions can cost up to $74.99 per month.
But understanding exactly how they get to their scores is not easy. Detectors do not hand out all of their technical workings, including which texts were used in training. Nor do they give users a precise breakdown of the score.
Broadly, what the detectors look for are patterns: not one supposed tell-tale AI word, like the frequently cited “delve”, but combinations that match the decisions a large language model might make.
There are different ways to do this. Researchers at the University of California, Santa Barbara came up with “Divergent N-Gram Analysis” (DNA-GPT) in which a text is chopped in half and AI is used to write a new ending — after which the two halves are compared.
Pangram, which came out at the top of a recent study of detector reliability by researchers at the Vrije Universiteit Brussel, took a more time-consuming approach. It gathered millions of human texts written before ChatGPT was released — everything from Amazon reviews to poetry to long-form essays — and then generated “synthetic mirror” texts from various chatbots and fed the pairs into its model in order to train it to tell them apart.
The training is updated as new AI models appear. “It took us a year to get better than the rest of the market,” says Spero. “And another year working on the problem of AI-assisted versus fully AI text. And now we have a good model.”
GPTZero, one of the earliest detectors created by 26-year-old Edward Tian, who once worked as an investigative researcher for Bellingcat, compares the patterns AI detectors look for to a mosaic. His software, which was used recently to identify AI hallucinations in reports compiled by EY and PwC, also uses a classifier trained on AI and human documents.
Text uploaded for AI detection is broken down into fragments and processed as tokens that are converted into numerical representations. These are compared to numbers generated by AI and human text to see if there are overlaps — not only in word choice but word placement, frequency, sentences and paragraphs. The more text there is, the more patterns there are to look at.
Features of text written by AI
• AI writing features the rule of three (X, Y and Z)
• Em dashes with a list of words afterwards
• ‘Can X but can’t Y’ formulation
• ‘Deeper’ is one of the words often seen in AI writing
Text edited by a human
• Repetition of ‘from’ changes the rhythm and creates uneven lines
• The long run-on sentence is broken up
Use of a humaniser
• Writing is less polished
• Same rule of three appears
• ‘Can X but can’t Y’ formulation remains
• Rewritten final part is a rambling, run-on sentence
This is why the leading detectors can still spot AI use even after someone has changed a few words or rearranged some lines. Some have also added plagiarism analysis and AI image identifiers.
The founders of detection companies know that they are in a race with AI labs, who have an interest in making their output as natural as possible, and “humaniser” tools designed to scramble and rewrite text to remove obvious signs of AI use.
The real enemy, however, is false positives — the risk that a model might incorrectly flag human text as AI generated. OpenAI removed its own detection programme in 2023 after reporting a 9 per cent false-positive rate. In 2024, one user claimed that a detector misidentified the Declaration of Independence as A widely reported Stanford study found that text from non-native English writers was more likely to be wrongly flagged as AI, something the authors suggested might be due to their more limited linguistic variability and word choices.
This is where an AI detector’s threshold for evidence becomes important.
Economist Brian Jabarian at Carnegie Mellon University published research last September that compared three commercial AI detectors: Pangram, OriginalityAI and GPTZero. He fed in nearly 2,000 human-written text samples alongside AI-generated equivalents and then set various evidence thresholds. Dial the threshold down, and while more text is identified as AI-generated, the chance of a false positive rate slightly rises.
Pangram says its latest model has a false positive rate of just 0.0041 per cent, or roughly one false positive for every 24,000 documents. The question is whether we are willing to accept the risk of that one false accusation. “That’s a decision that isn’t up to tech companies,” says Jabarian. “It’s social.”
Nazir claims that he is one of those who have been falsely accused. After his short story won the Caribbean regional Commonwealth Prize in May, an anonymous X account posted a screenshot that showed Pangram identified it as 100 per cent AI-generated.
Others chimed in to say they knew it was AI even without the detector. “I could tell it was GPT in about 5 seconds”, wrote Nabeel S Qureshi, a former Palantir engineer.
Commentators pointed to what they considered AI giveaways. “Not the bees’ neat industry . . . but a belly sound” was a “not x but y” formulation common in AI output. The line “she had the kind of walking that made benches become men” was seen to be a nonsensical AI metaphor, while “the quiet chores, the patient hands, the unlit lamp” was a rule-of-three typical of AI text. A few people began to wonder whether the writer himself was real — pointing to his strangely symmetrical photograph.
Speaking from his sunny, plant-filled home in Trinidad, Nazir says that yes, that picture was digitally altered to smarten it up but no, his story was not written with AI. The line about benches, for example, that was his way of saying that a woman could be so beautiful that inanimate objects would come alive.
“It’s a literary technique,” he says. “Don’t they understand mysticism, these readers? There’s a line in one of Salman Rushdie’s books where people float through the air — that didn’t happen either.”
Nazir, a 62-year-old retired civil servant, is scruffier than his infamous photo, with salt-and-pepper stubble and thick black glasses that he rests on his forehead. Describing the events of the past few months, he switches between good-natured amusement (“that picture was looking really good”), hurt at the critique of his writing and exasperation at the claims he used AI.
The story he entered was, he says, plucked from a childhood memory of walking to school through cane fields. His writing style, which some readers described as stilted and artificial, is the result of a career spent writing technical reports and a lifetime reading poetry by the likes of India’s Rabindranath Tagore and the Palestinian poet Mahmoud Darwish. But not AI — not even for editing. “I don’t use AI, I use GI,” he says. “God’s intelligence.”
As his health is poor — neuropathy makes it difficult to type and he is undergoing chemotherapy for lymphoma — he dictates his poems and stories using a speech-to-text tool. But, he repeats, he doesn’t use any sort of AI writing tool.
His critics are unmoved. “It was pretty obvious,” says Ethan Mollick, an AI researcher and associate professor at the Wharton School of the University of Pennsylvania, who criticised the story online. “AI is good at a lot of things but writing fiction is not one of those things.”
Tuhin Chakrabarty, an assistant professor in the computer science department of Stony Brook University SUNY, who co-authored the paper on AI use in ebooks, points out that the story contains six instances of the word hum — something he says ChatGPT is known to use in fiction. He has no issue calling out individuals he thinks are trying to pass off AI writing as their own. “I write on copyright law and I think there should be stigma [about using AI],” he says.
The Commonwealth Foundation conducted an investigation and stood behind Nazir. But Granta, the literary magazine, announced that it would no longer publish the winning stories. “At this stage there is really no way of determining whether a piece is authentic or not,” says Sigrid Rausing, Granta’s publisher. Still, she adds, she is not opposed to AI. “If a writer creates an experimental text with AI material and is transparent about it, I would be happy to read it and publish it, if it’s good.”
Such experiments seem unlikely. In 2023, author James Frey told French magazine Pompidou+ that he was using AI in the composition of his new book, saying that he would not disclose which parts were written by him and which by AI. This playful text is yet to materialise.
The same silence hangs over commercial work. Despite speaking openly about AI and its positive effects, many companies seem reluctant to clarify their exact use of it in written reports. When AI usage leads to errors, as it has for work done by KPMG, EY, PwC and Deloitte, companies have been criticised.
“If these were human errors, would the accusations be as loud?” asks Ashley Williams, a lawyer and head of technology at Mishcon de Reya in London. “I think this shows that we are all still working out how we should use AI — where the line should be drawn.”
One consultant said transparency was lacking because companies worried that clients would regard AI-generated reports as intrinsically lower quality. “[AI] is very good at producing the parts that are formulaic, but clients may not realise that you still need a professional to review and rewrite documents. There is a stigma that AI-generated text carries less weight.”
At the end of this year, the EU plans to introduce watermarks to AI text, something campaigners have been demanding for years. Anthropic has already stepped forward to say that its new Claude models will contain these imperceptible marks.
Yet instead of clearing things up, watermarks could complicate them further. Editing, paraphrasing and copying text could erase the marks. Anthropic points out that the absence of watermarks will not prove that AI was not used.
Mollick, at the Wharton School of the University of Pennsylvania, thinks the fact that no detector or watermark is 100 per cent foolproof means we have already passed AI’s “K-T boundary” — the geological marker of an extinction event.
“Authenticity is going to be a problem everywhere,” he says. “We’re still in the middle of negotiating what parts of writing we want to keep human and which parts we accept can be AI.”
Codes of conduct are in flux. If AI is unacceptable when writing literature, is it fair game for more prosaic text like legal drafts and audits so long as the quality is high and the time saved is reflected in the price? Are AI-drafted emails a useful tool or an insult to those on the receiving end? Must AI drafts and edits always be disclosed? And will accusations of AI based on intuition and detectors lead to greater surveillance and semantic contortions?
“I think the future will be an equilibrium — a balance of human and AI writing,” says Tian, creator of AI detector GPTZero. He believes accusations of plagiarism will die down as more people integrate AI into their lives. But, he says, we still need to find ways to communicate with one another without machines mediating.
“Writing is more than just output. It reflects our critical thinking,” he says. “There is something beautiful about it that is worth preserving.”
*The FT editorial code of practice states that generative AI tools must not be used to write an article or article text for publication, nor to create other reader-facing content.