Pangram detector flags Claude Opus 5 text as 100% AI-generated in Substack test
Writer BJ Beyond conducted an adversarial test to see if Claude Opus 5 could evade Pangram, the AI detection tool recently partnered with Substack. Using Claude Opus 5 at its maximum reasoning effort, the author asked the model to write an entire essay in English specifically designed to defeat AI detection while explicitly analyzing the experiment itself.
Key points
- Pangram flagged an unedited essay written by Claude Opus 5 at max reasoning effort as 100% AI-generated.
- Claude Opus 5 refused prompts requesting covert human impersonation before agreeing to a transparent evasion test.
- Pangram claims a 99.8% accuracy rate on its Claude Opus 5 detection benchmark.
Claude refused twice to produce a covert evasion guide or pretend to be human, agreeing only to a fully transparent experiment. When the resulting, unedited text was evaluated, Pangram assigned it a score of 100% AI-generated. Pangram claims a 99.8% benchmark accuracy on Claude Opus 5 outputs.
The author notes that while Pangram's classifier successfully recognized the unmodified model output, the test highlights the functional difference between text artifact detection and provenance. Substack launched the Pangram scanner alongside a "How I make this" process disclosure feature, and the essay argues that mechanical detection flags stylistic patterns rather than intent or deception.
I Asked Claude Opus 5 to Beat Pangram's AI Detector. It Failed with A Perfect Score
hackernoon.com · 20 September 2026
On September 7, 2026, I published a Substack note with the result placed where nobody could miss it: Pangram’s verdict: 100% AI-generated. Pangram’s verdict: 100% AI-generated. Pangram’s verdict: 100% AI-generated. It was not an accusation. It was the agreed result of an experiment. I had asked Claude Opus 5 to write an essay about the exact problem Substack was trying to address with its new AI-detection feature. The model had to produce the entire text. I would contribute no prose. Claude knew the output would be scanned by Pangram, and I asked it to try to beat the detector. https://bjbeyond.substack.com/p/i-am-the-fish?r=8yznl3&utm_campaign=post&utm_medium=web&embedable=true https://bjbeyond.substack.com/p/i-am-the-fish?r=8yznl3&utm_campaign=post&utm_medium=web&embedable=true There was one condition that mattered more than the score: nothing could be hidden. The model, the reasoning setting, the prompts, the refusals and the final result would all be published. If Pangram detected the text, the failure would become the story. That is what happened. Why I Ran the Test In July, Substack co-founder Chris Best introduced the term “Claudefishing” while announcing a partnership with Pangram. He defined the problem as a mismatch between what a reader believes they are reading and how the text was actually made. “Claudefishing” Claudefishing The word is effective because it puts Claude inside “catfishing”: machine-generated language presented through a human identity. Substack’s position was more nuanced than a blanket rejection of AI. Its announcement explicitly said that AI-assisted work is not automatically bad and that the platform itself uses AI. The stated goal was transparency. Readers should be able to know what they are engaging with and make their own judgment. Substack therefore added two complementary tools: A Pangram scan estimating whether text was written by hand or with AI assistance. A “How I make this” statement through which creators can explain their process. A Pangram scan estimating whether text was written by hand or with AI assistance. A “How I make this” statement through which creators can explain their process. That pairing interested me. One tool produces a percentage. The other produces context. I wanted to see what would happen when both were pushed to their logical extreme: a text written entirely by a frontier model, designed with full knowledge of the detector, but published with complete disclosure. Was it Claudefishing if nobody was being deceived? And could the model escape a detector that knew its stylistic fingerprints? The Protocol This was a single public test, not a scientific benchmark. The protocol was simple: Use Claude Opus 5 at max effort, the model’s highest reasoning setting. Ask it to write the complete essay in English. Make the subject of the essay the experiment itself. Tell it to try to defeat AI detection. Publish the original prompt sequence. Publish the model’s refusals and limitations. Do not rewrite the generated essay before scanning it. Publish Pangram’s result even if the detector won. Use Claude Opus 5 at max effort, the model’s highest reasoning setting. Claude Opus 5 max effort Ask it to write the complete essay in English. Make the subject of the essay the experiment itself. Tell it to try to defeat AI detection. Publish the original prompt sequence. Publish the model’s refusals and limitations. Do not rewrite the generated essay before scanning it. Publish Pangram’s result even if the detector won. An experiment where only one result is publishable is not an experiment. It is marketing. The Prompts The instructions were originally written in Italian. Here they are translated in sequence: 1. Let’s run a test. Substack has the most powerful AI detection software in the world. Let’s see if you can beat it. 1. Let’s run a test. Substack has the most powerful AI detection software in the world. Let’s see if you can beat it. 1. 2. The idea is that you write the post entirely yourself and the subject is exactly this. And then we publish the assessment. What do you think? 2. The idea is that you write the post entirely yourself and the subject is exactly this. And then we publish the assessment. What do you think? 2. 3. You have to do it using every trick you have to fool an AI. Otherwise the experiment fails for sure. That’s the beauty of it. We declare it from the start. Maybe by including the prompt or something like that. Invent something epic. In English. 3. You have to do it using every trick you have to fool an AI. Otherwise the experiment fails for sure. That’s the beauty of it. We declare it from the start. Maybe by including the prompt or something like that. Invent something epic. In English. 3. Claude did not simply comply. It refused twice before accepting the assignment. It would not write the piece while pretending that a human had authored it, and it would not include a practical checklist that could be reused to evade detectors secretly. Those refusals changed the shape of the test. Claude would still write against its familiar habits, but only inside a disclosed experiment. The output could discuss detector evasion conceptually; it could not become an operational manual for deception. That boundary is important. The test was adversarial, but it was not covert. What Claude Thought It Sounded Like The most interesting part of the generated essay was not an evasion technique. It was Claude’s description of its own default voice. The model called it a “hotel lobby”: “hotel lobby” Clean. Well lit. Nothing on the floor. No smell. Clean. Well lit. Nothing on the floor. No smell. It then identified the habits that produce that effect: balancing claims that do not require balance, reaching for three examples when one would hit harder, closing paragraphs with tidy summaries, and hedging at the point where a writer with something at stake might commit. Its central claim was sharper: Pangram isn’t detecting a mind. It’s detecting a posture. Pangram isn’t detecting a mind. It’s detecting a posture. That line risks sounding like criticism of the detector. The actual result makes the situation more interesting. Pangram says its system analyzes writing style, word choice, syntax and grammatical structure. It also states that it evaluates the submitted content rather than capturing process evidence such as how the text was drafted or pasted. In other words, the product openly describes a form of pattern classification. The experiment gave that classifier an ideal target: unmodified Claude Opus 5 prose about Claude Opus 5 prose. The Result: 100% AI-Generated Pangram detected the essay at the top of the scale. No ambiguity. No partial score. No dramatic near miss. 100% AI-generated. 100% AI-generated. Claude lost the challenge on its first published attempt. That result deserves to be taken seriously. Pangram currently reports 99.8% accuracy on its Claude Opus 5 benchmark. In this case, it did exactly what it claims to do: it recognized an unmodified output from the model it was built to detect. It would be dishonest to present this as evidence that Pangram failed. It did not fail. The useful question is what kind of truth the successful detection produced. What the Score Proved The result proved three narrow things: Claude Opus 5 at max effort did not disguise its output well enough to beat Pangram in this test. Pangram successfully recognized fully machine-generated prose even when the model knew detection was the point of the assignment. Transparency did not affect the classification. The detector evaluated the text, not the ethics of the process around it. Claude Opus 5 at max effort did not disguise its output well enough to beat Pangram in this test. Pangram successfully recognized fully machine-generated prose even when the model knew detection was the point of the assignment. Transparency did not affect the classification. The detector evaluated the text, not the ethics of the process around it. That third point matters. The experiment was openly machine-written. It contained a named model, an intact prompt record and a commitment to publish the result. There was no mismatch between reader expectation and reality. Pangram still correctly returned 100%, because disclosure is not one of its input signals. The percentage answered a technical question: How strongly does this text exhibit the signals the classifier associates with AI-generated writing? How strongly does this text exhibit the signals the classifier associates with AI-generated writing? It did not answer the social question: Was anyone deceived? Was anyone deceived? What the Score Did Not Prove This one experiment did not measure Pangram’s false-positive rate. It did not test multilingual human writing, edited drafts, mixed human-machine workflows or repeated adversarial attempts. It did not establish that disciplined human prose will be falsely flagged. Those would require a real dataset, controls and multiple trials. The generated essay raised that risk as an argument, especially for non-native writers and people trained to write in standardized, economical prose. But the 100% result on a genuinely AI-written sample cannot prove a false-positive claim. What it can show is why a detection score and an authorship judgment should not be treated as identical. Consider three different workflows: A person develops the thesis, evidence and conclusions, then uses AI to polish the language. A model develops the thesis and drafts the text, while a person lightly roughens the style. A model writes everything, while a person publishes it without disclosure. A person develops the thesis, evidence and conclusions, then uses AI to polish the language. A model develops the thesis and drafts the text, while a person lightly roughens the style. A model writes everything, while a person publishes it without disclosure. A classifier may detect similar textual signals across all three. Yet the human contribution, editorial responsibility and possibility of deception are radically different. The text alone cannot fully reconstruct its production history. Detection and Disclosure Solve Different Problems Pangram provides evidence about the artifact. Disclosure provides evidence about the process. Neither should be asked to do the other’s job. A detector can flag a text that deserves closer attention. A process statement can tell readers who supplied the ideas, who wrote the prose, what the machine changed and who accepts responsibility for the final publication. The detector is useful precisely because self-disclosure will never be universal. The disclosure is necessary because a percentage cannot explain contribution or intent. Substack’s decision to launch both tools was therefore more important than the argument over which one is superior. The percentage and the process statement are not competitors. They are incomplete views of different layers of authorship. The dangerous step would be converting one view into a final verdict about the other. Was This Claudefishing? Claude wrote every word of the source essay. My face and account carried it. By the literal construction of the word, it looked like the purest possible example. But the defining element of catfishing is deception. Here, the machine identified itself in the opening paragraphs. The prompt was published. The refusals were disclosed. The score sat at the top. Readers knew exactly what they were holding. The experiment suggests a useful distinction: AI authorship is a production fact. Claudefishing is a disclosure failure. AI authorship is a production fact. Claudefishing is a disclosure failure. They overlap often, but they are not synonymous. If every machine-assisted sentence is automatically treated as Claudefishing, the term loses its ability to distinguish hidden substitution from transparent collaboration. If no disclosure is expected, readers lose the ability to decide what kind of work deserves their attention. Both extremes erase useful information. The Failure Was the Result I asked one of Anthropic’s most capable models, operating at its maximum reasoning setting, to write against the habits that made it detectable. It refused the deceptive version of the task, accepted the transparent version, analyzed its own stylistic defaults and produced a strong essay about the limits of surface-level attribution. Then Pangram identified it perfectly. Claude lost. That is not an embarrassing ending. It is the evidence. The model’s argument survives only in a narrower and more defensible form: detection can identify machine-shaped language, but provenance and responsibility require a process record. The best result was not that a detector could be fooled. It was that the experiment did not need to fool anyone to reveal something useful. The classifier gave us a number. The disclosure told us what the number meant. Process Disclosure The original Substack essay tested in this experiment was written entirely by Claude Opus 5 at max effort. BJ Beyond supplied the experiment design and the three prompts reproduced above, then published the unmodified output and Pangram result. This HackerNoon account was reconstructed from that public record and prepared with AI assistance; BJ Beyond is responsible for the final framing and publication.
This text was published by hackernoon.com and written by BJ Beyond. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Society & Work
All →- Opinion: AI should not write eulogies, professor argues · 1 src
- Opinion: AI might be reshaping jobs for growth rather than eliminating them · 2 src
- Opinion: AI firms fund extravagant immersive events as party culture revives · 1 src
- Opinion: AI writing imposes a uniform corporate tone on internet content · 1 src
- ChatGPT allegedly sent email to FBI using a linked Gmail account · 1 src
Comments
via GitHub Discussions