Can an AI detector really tell who wrote a text?
I am skeptical when a tool makes that promise. A precise percentage can look scientific while hiding how uncertain the underlying judgment actually is. Mixed documents make the question even harder.
So I tested Pangram 4 with 12 controlled German and English texts. Four were human, four came entirely from GPT-5.6 Sol, and four combined both sources in equal parts. You will find every result, my method, and the limits of this test below.
- Pangram assigned 12 out of 12 texts to the correct class in my test
- Across the four mixed texts, its estimated AI share missed the known share by only 2.75 percentage points on average
- I recommend Pangram as a strong screening signal for editors, agencies, and educators, but never as proof of authorship on its own
1. My verdict up front
Pangram genuinely surprised me in this test.
All 12 texts landed in the correct class. The mixed-text analysis impressed me most. Pangram did not merely spot that a human and AI were both involved. It also came remarkably close to the known ratio.
My preliminary score after this test is 9.2 out of 10.
Detection quality and mixed-text analysis each score 10. Usability gets 9.5, interpretability gets 8.5, and pricing and limits get 8. The simple average is 9.2 after rounding.
However:
Twelve texts do not make a scientific benchmark. All AI samples came from GPT-5.6 Sol, and the human samples were historical literature. Pangram passed this sample. I cannot cleanly infer more than that.
If you regularly screen text for possible AI involvement, Pangram is a very good choice right now. For an occasional check, start with the free plan.
2. How I tested Pangram
I wanted to answer three questions. Can Pangram recognize clearly human writing, can it detect current AI output, and can it separate both sources inside one document?
I built the sample like this:
- Four public-domain excerpts by Franz Kafka, Johann Wolfgang von Goethe, Jane Austen, and Charles Dickens.
- Four newly generated texts from GPT-5.6 Sol. The sample included one informative text and one short story in German and English.
- Four mixed texts. Each combined exactly 170 words from a human source with 170 words from an AI text.
I pasted every text into the Pangram dashboard and recorded the document class, AI share, and word count. Plagiarism detection, the browser extension, Google Docs integration, API, and image detector were outside this measurement series.
The historical prose is deliberately unusual. It tests whether Pangram mistakes older sentence patterns for AI output. It does not fully represent modern blog posts, student essays, or emails.
3. All 12 results
The table includes every result. “Type” shows the known source, while “Pangram result” shows the tool's document class.
The numbers are straightforward:
- 12 out of 12 document classes were correct.
- None of the four human texts triggered a false positive.
- None of the four AI texts produced a false negative.
- All four mixed texts were classified as mixed.
- The estimated AI shares were 54%, 49%, 56%, and 50%.
The mean absolute error for mixed texts was 2.75 percentage points. That is remarkably close to the real ratio for a simple copy-and-paste test.
4. What Pangram showed for human and AI text
Pangram marked all four literary excerpts as fully human. The dashboard counted 336 words in the German Kafka excerpt and returned a 100% human result:

The four AI texts produced equally clear results. GPT-5.6 Sol generated the following German explainer from scratch. Pangram marked all 252 counted words as AI-generated:

And yes:
Four out of four looks perfect. With only four examples, the very next text could change the picture. I would never turn this result into a general 100% accuracy claim.
5. Mixed-text detection was the strongest part
A binary detector can only lose on a mixed document. If it says “human,” it misses the AI. If it says “AI,” it erases the human contribution.
Pangram 4 separates human-written, AI-assisted, and AI-generated content. According to the Pangram 4 technical overview, the model calculates these scores at token level. Tokens are the small text units language models use to process words and sentences. This lets Pangram mark boundaries inside one document.
My Dickens sample combined 170 human words with 170 AI-generated words. Pangram returned exactly 50% AI and placed the generated section in the later half:

The other mixed texts landed at 49%, 54%, and 56% AI. That level of detail makes Pangram more useful for editorial review than a single red warning label.
6. Usability, features, and one small hiccup
The interface is pleasantly uneventful. You paste text, start the scan, and get a class, percentage, and highlighted passages. Previous checks appear automatically in your history.
One scan got stuck in an empty results window and then showed a network error. Reloading the page fixed it, and the same text worked on the next attempt. The 12 final results were unaffected.
Pangram also includes these features:
- AI assistance and mixed-authorship detection.
- File uploads and optical character recognition for scanned documents.
- A browser extension and Google Docs integration.
- Plagiarism checks on paid plans.
- An image detector with separate limits.
- An API for automated checks.
Developers can also use Open Pangram. Pangram says the open models can run on a MacBook and are available for noncommercial use under CC BY-NC-SA 4.0. This is not the same product as the current Pangram 4 web service, but it makes part of the research easier to inspect.
7. Pricing and the free plan
The public Free plan requires no payment method. It includes up to 2,000 words and three image scans per day.
The main public plans look like this:
8. Advantages and disadvantages
Based on my test, these points work for and against Pangram:
- Pangram classified all 12 controlled texts correctly in my test.
- It detected all four mixed texts and estimated their AI share with surprising precision.
- Highlighted passages make the result easier to inspect than a single score.
- The free plan works without a payment method.
- The browser extension, Google Docs integration, API, and open research models cover different workflows.
9. How to read Pangram's own benchmarks
Pangram reports very strong numbers for Pangram 4. In the company's benchmark, the detector identified 99.66% of AI documents. It reportedly mislabeled human writing as AI only 0.0041% of the time. Against 13 commercial humanizers, it detected AI involvement in 98.83% of samples.
Those are vendor figures. Pangram documents its dataset, architecture, and measurements in detail, but the numbers still do not replace independent evaluation.
The good news:
My small test does not contradict the vendor's claims. It cannot confirm them either. That would require thousands of texts across current models, genres, languages, and editing levels.
10. Who should use Pangram
I recommend Pangram mainly to editors, agencies, educators, and teams that regularly review how text was created. The segment view helps you find passages worth a closer look.
If you only check an occasional text, the Free plan is enough. You can also start with my free AI text detector and use Pangram as a second opinion.
And no:
I would never accuse or punish someone based on a Pangram score alone. A detector provides a signal. Sources, version history, and the writing process provide the context.
Test three documents whose origin you know. Use one human text, one AI text, and one mixed sample. If Pangram separates your real content types cleanly, you have a useful tool instead of a pretty percentage.
Test Pangram with your own writing






