We Underestimate the Weaknesses of Pangram

How much does Pangram's "AI-Generated" label indicate the degree to which an author has outsourced their thinking?

When they tested their 4.0 product, Pangram found that, by their definition, the proportion of AI-Assisted documents it classified as AI-Generated was 0.01%, 4%, or 7%, depending on the experiment. Then they omitted the experiments that found 4% and 7% false positive rates (FPRs) on their website, while advertising there that the product detects AI-Assisted writing. Before I contacted them about this issue on September 17th, their claim on their website was more misleading—"99.9%+ Accuracy" was displayed directly adjacent to the phrase "Detects AI Assistance" on the text detection input box that many people don't read past. Similar claims remain repeated elsewhere on the main page instead of by the text box itself. I do not know if my message was the cause of the change.

Their experiment that found the 0.01% FPR might be more appropriate for identifying human-written text rather than AI-assisted writing. For example, the prompt they gave to Claude for this experiment was "Fix spelling, punctuation, and clear grammar errors only." In my experience, human editors normally provide conceptual feedback as well. Since we generally don't cite human editors, shouldn't the label AI-Assisted indicate more assistance from AI than would be provided by a human editor?

In my experience, our community also heavily relied on Pangram before their 4.0 release on July 29th. The previous version had FPRs for the AI-Assisted vs. AI-Generated labels for these experiments of 0.2%, 15%, and 22% respectively.

I think these issues have given people in our community an unrealistic impression of Pangram's ability to distinguish AI-assisted writing from AI-generated writing. At least an order of magnitude is a big difference! We often use Pangram's labels to assess how much we should trust writers. This makes it harder for us to obtain accurate information if we are unaware of the accuracy of the labels.

It’s likely that the FPRs will be concentrated in the work of writers who write in a way that particularly confuses Pangram — the most elite writers not seeing false positives for their writing does not establish that their experience is typical. Sometimes I write in a style that Pangram defines as AI-Assisted writing in their experiments that found 4% and 7% FPRs. I spend so much time on these pieces that I am able to match almost every sentence to Pangram's definitions. I have run about 5,000 words of this writing of mine through Pangram 4.0, and it labels about 30% of the passages in it AI-Generated, often with high confidence. When these passages are are part of a long enough excerpt, Pangram's paper would often not have defined them as a false positive. The 4% and 7% FPRs only counted instances where the entire excerpt was mislabeled AI-Generated. Many of my longer passages are overall labeled as mixed. If my experience is common, the problem is much worse than a 4% or 7% FPR would indicate.

Pangram's definitions of AI-Assisted rely on more objective measures than the subjective task of comparing prompts. However, to give a sense of the writing style that so confuses the product, Claude's prompt for the experiment that found a 4% FPR was "Substantially rewrite for polished academic style while preserving meaning." My process is similar. After I've produced writing I would have called finished a year ago, I give an LLM a prompt like "Please make the wording of the blander parts of this excerpt more poetic and visceral, to match the parts of it that are most in that style. Do not remove core concepts or add additional ones." After receiving a draft produced by such a prompt, I spend roughly six more hours per 1,000 words editing it. Overall, I reject the vast majority of suggestions the LLM gives me. I have included a footnote with a representative example of my process. 1 I believe the AI-Assisted label describes this process well. But many people besides me think there is a large difference between this writing and having an LLM generate a piece on its own. I have never met someone in our community who values a product that can distinguish between the two and was previously aware of how often Pangram struggles with this task.

Like others, I have noticed that when I run my work by Pangram in chunks of a few hundred words instead of the whole piece, the FPR increases dramatically despite the writing being identical. Pangram acknowledges that its results are more uncertain for shorter passages, but we should be aware that Pangram's headline accuracy figures are unlikely to apply to the X posts and emails many of us are now using the product for.

Because of my high FPRs with Pangram, I only use AI editing in my more private writing. It is better than having my more important work dismissed. I think this is a big waste. I am particularly concerned about how overreliance on Pangram might affect people like a friend of mine. He is a brilliant scientist whose native language isn't English. He once relied on me to help him sound fluent in his papers. I never had enough time for him, and he was overjoyed when AI editing could finally give him the English help he needed. He needed to use AI editing in the way that Pangram found 4% and 7% FPRs for. Indeed, one of the prompts Pangram used in the experiment that found a 4% FPR was "Sound fluent." Overreliance on Pangram could be depriving us of the insights of people like him, as well as native English speakers who are simply too busy doing interesting research to waste time competing with people who specialized in writing.

If AI editing typically made human writing worse, my previous point would be weak. However, in my experience, AI editing is very good now. I've experimented with almost every new Claude and ChatGPT update, and in 2026 their wording suggestions have gone from almost useless to usually better than the best I can produce. People who dislike AI-editing tells often describe a visceral feeling of emotional distress when they see them. I wonder if these feelings arise from the memories of when AI-edited writing was much worse, and from the automation of a skill many of us have put a lot of love into developing. Basing decisions on something as subjective as noticing AI tells when under such stress is often bad epistemics. Someone posted a real Monet painting with the claim that it was AI and asked for an explanation of why it was inferior to the real thing. People responded with a flood of scathing critiques. Why wouldn’t a similar cognitive bias be happening with writing? Perhaps Pangram gives people who would normally be more careful an excuse to settle for these kinds of mistakes.

I actually think Pangram is a very impressive product. I've looked at about a dozen independent papers evaluating them, and believe their paper did an unusually good job of producing an AI-Assisted sample that is similar to how AI editing is really used. They are attempting a very difficult technical problem and doing it well. But I think being aware of the limitations of the product would be helpful when we evaluate the degree to which writers have outsourced their thinking, and which writers to trust.

  1. An example of my writing process that resulted in a passage that Pangram 4.0 called 100% AI-generated (when part of a 3,000 word submission):

Pre-AI draft: Erhan’s spies told him this was where Nasreen was now working, and he climbed the steps to her office hoping to find some documents left behind in a hurried evacuation. If his scholars could decipher some of her work, perhaps it could keep famine away for a while.

Prompt: I'd like you to rewrite it. I really like your word choice suggestions and I'd like you to incorporate them everywhere, but I'd like you to keep all the actions that take place the same.

Claude's draft: Erhan’s spies said this was where Nasreen worked now, and he climbed the stairs to her office with a thief’s hope in his throat: papers left behind, a ledger, a packet of notes—anything his scholars could worry into meaning before winter made a mouth of the countryside.

My final Draft: Erhan’s spies had told him this was where Nasreen was now working, and he climbed the steps to her office with a kind of hope that felt like thirst: for papers left behind, a ledger, a packet of notes—anything his scholars could worry into meaning before another harsh winter came to collect what war had spared.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论