Skip to main content

16 AI Text Detectors Compared: Tests, Prices and Limits

Compare 16 AI text detectors using documented tests, screenshots, price checks, and limits. See where the strongest signals still fall short in real use.

FHFinn Hillebrandt
AI Tools
16 AI Text Detectors Compared: Tests, Prices and Limits
Links marked with * are affiliate links. If a purchase is made through such links, we receive a commission.

An AI detector can sound far more certain than it is.

A percentage, a colored passage, and a label such as AI-generated can be useful signals. They can’t establish authorship on their own. That matters when you edit client work, review student submissions, or decide whether to trust a manuscript.

I compared 16 AI text detectors and research demos. The article separates documented hands-on tests from provider claims and products that were unavailable when checked. You’ll see the full test results, screenshots, price information, and the limits that matter before you rely on a score.

TL;DRKey Takeaways
  • Pangram matched all four pure full-text controls and labeled both mixed controls as mixed. Its section boundary was still imperfect in one English mixed text.
  • Originality.ai, Copyleaks, Sapling, and the retired German Gradually detector each produced documented false positives or missed detections in the available series.
  • Use a known human text that resembles your own work before you trust any detector with a customer document or a high-stakes decision.

1. What this comparison can and cannot tell you

There is no universal accuracy ranking in this article.

Six detectors received documented text checks. Pangram and the retired German Gradually detector processed the same six full texts. Sapling and Gradually also received six shorter controls. Originality.ai, GPTZero, and Copyleaks have older, limited controls that remain useful as observations, not as a leaderboard.

The other ten entries cover access checks, public product pages, discontinued routes, or research demos. Where a text test did not run, I don’t turn marketing claims into a performance judgment.

Each active product has a screenshot directly below its documented result or access information. Screenshots labelled archival show a former interface. Their visible figures are not current test results.

2. How the practical tests were set up

Six correctly labeled texts are not a reliable accuracy rate.

The most recent complete Pangram series ran on September 19, 2026. It used archived Wikipedia revisions from 2021 or earlier as known human references, fresh AI-generated explainers, and prebuilt mixed documents. The German topic was photosynthesis. The English topic was the water cycle.

The generated counterparts came from the configured GPT-5.6 Luna route. They were written without a human source text and were not revised after detector results. The mixed documents preserved the original passages, so their sources were known. They do not recreate a normal editorial collaboration.

Same full texts, separate claims

Pangram and the retired German Gradually detector received the same six inputs. The table reports document classes. Their percentages use different methods and cannot be compared as if they measured the same thing.

Pangram ran with the visible 4.0 model in an existing licensed account.

Source: own control runs, September 19, 2026. Word counts use whitespace. Mixed documents reuse their source texts.

Known sourceGerman, Human
Words406
PangramHuman
Retired German Gradually detectorAI (false positive)
Known sourceGerman, AI
Words506
PangramAI
Retired German Gradually detectorAI
Known sourceGerman, Mixed
Words451
PangramMixed
Retired German Gradually detectorAI, no mixed class
Known sourceEnglish, Human
Words578
PangramHuman
Retired German Gradually detectorAI (false positive)
Known sourceEnglish, AI
Words496
PangramAI
Retired German Gradually detectorAI
Known sourceEnglish, Mixed
Words472
PangramMixed
Retired German Gradually detectorAI, no mixed class

Four pure texts in direct comparison

This chart includes only the two human and two fully AI-generated texts. Both services received the same inputs. The bars show matching document classes in four cases, not a general accuracy rate.

Pangram
4 of 4
Gradually
2 of 4
Own control runs, September 19, 2026
gradually.ai

The two mixed documents are excluded because they reuse the same source texts. Pangram labeled both as mixed. Gradually did not offer a separate mixed class and labeled both as AI.

Sapling does not belong in that full-text table. Its responses included only the first 2,000 characters. In mixed texts, that could exclude the appended AI section. Sapling received a separate short-text series instead.

What these results do not establish

The series does not cover every text type or current writing model. Each language lacks independent repetitions on different subjects. Mixed documents do not enlarge the independent sample when they reuse the same inputs.

A prepared text counts only after the service actually analyzes it. A missing response is not a false detection. Older literary controls also do not become part of one overall rate with current factual texts.

The older series lacks the original prompts and provider responses. Its detector screens are preserved, but the provenance is less reproducible than the newer factual controls. I therefore do not use those older cases to rank tools.

3. Tested AI text detectors

A readable report only helps if it still belongs to the text you checked.

The sections below describe real controls and selected usability checks. They are not complete audits of every plan, integration, or data practice. Public screenshots were captured on September 20, 2026. Visible marketing claims belong to the provider, not to these test results.

3.1 Pangram found the mixed documents, but not every boundary

Pangram returns a document class and highlighted passages. That makes it possible to inspect the difference between a fully AI-generated text and an inserted section.

The mixed reports showed 24% AI for the German document and 22% for the English document. According to the Pangram 4 model card, these shares are weighted by characters. They are not directly comparable with the word share in the inputs.

One issue was visible without interpreting any percentage. In the English mixed text, the first 19-word AI sentence appeared inside the human-marked section. The overall Mixed class still fit the document.

This is Pangram’s public input interface, not the report from that mixed test:

Pangram’s public interface with an empty text field and an image-detection switch

History worked in the usability check. All six reports could be reopened. Earlier document checks also included importing and analyzing a DOCX file and a text-based PDF. TXT upload was not available. OCR, batch processing, and team functions were not tested.

Pangram’s highlights can help you inspect an unusual passage. Do not assume that every colored boundary is exact. My separate Pangram review covers the product in more detail.

The Pangram pricing page lists a free entry plan and paid word allowances. I used an existing license for these controls and did not buy or extend a plan.

See Pangram

3.2 Originality.ai combines checks and produced false positives

Originality.ai combines AI detection and plagiarism checking in one account. Its free Basic plan included AI detection only. Plagiarism checking remained off for these detection controls.

An older series from September 11 and 12 ran 12 controls through the multilingual model. It included literature excerpts, AI explainers, and mixed texts, six in German and six in English. The human references did not receive consistent labels.

Goethe received an Original class with 99% confidence. A human Kafka excerpt was labeled Likely AI with 97%. The same underlying error occurred with Dickens in English. Pure AI texts and mixed texts all received AI labels. That does not show whether Originality identified the exact mixed shares.

The public check page shows the text field, upload option, and available free scans. The image contains none of my findings:

Originality.ai’s public checker with an empty text field and three available free scans

A TXT import and analysis worked. DOCX and PDF files could also be read, but this control did not run a separate analysis for those formats. A synthetic report was reachable through its known share link without a login. That does not establish whether other reports are public or how this share setting was created.

Review the Originality.ai pricing alongside your privacy requirements. The documented false positives make automated acceptance or rejection of writers a poor use case.

See Originality.ai

3.3 GPTZero showed uncertainty, then failed to return a usable result

GPTZero distinguished between human, AI, and uncertain results in an older control series.

The German Kafka text received a 99% human label. The German mixed text showed 50% AI and 50% human, so the interface treated it as uncertain. In the English controls, Austen was labeled human. The pure AI text and the mixed text received their respective expected class. These six values came from an earlier authenticated Basic scan run.

That result could not be repeated with the new test text on September 19. Two normal attempts returned no analysis response that could be assigned to the submitted text. It remains open how GPTZero would have labeled it.

The screenshot shows GPTZero’s public interface with example buttons, not a successful repeat test:

GPTZero’s public interface with an empty text field and examples for different text types

The GPTZero website lists German as a supported language. Its student page also lists 10,000 words per month after sign-up. Check whether the service returns a reliable result in your own account. A good older run cannot replace that check.

See GPTZero

3.4 Copyleaks failed on a human control

Copyleaks accepted two German texts in the free check.

It labeled the human Kafka excerpt as 100% potentially AI-generated. The genuinely generated factual text also received 100%. The credit limit stopped the following mixed-text test.

That is one concrete false positive. It does not establish a false-positive rate for English, mixed documents, or the product as a whole.

This is the layout of Copyleaks’ public web checker. The empty field does not show a new test result:

Copyleaks web interface with an empty AI detector field and separate tabs for other checks

The Copyleaks product page lists a web checker, browser extension, and API. The extension and API were not tested here. The historical control used the free checker available at the time. It does not erase the observed false positive.

3.5 Sapling processed only part of long inputs

Sapling returns an AI score and sentence markings. Its interface can start an analysis automatically.

In the full-text experiment, the responses contained only the first 2,000 characters. That changes the question for mixed documents. When the AI passage sits at the end, a service may inspect only the human opening.

Sapling and the retired German Gradually detector therefore received the same six short inputs. They were fixed in advance and ranged from 1,434 to 1,573 characters. All six were processed in full.

Source: own short-text controls, September 19, 2026. Sapling values show one decimal place and are not text shares.

Short control textGerman, Human
Sapling AI score97.7% AI
Retired German Gradually classHuman
Short control textGerman, AI
Sapling AI score72.2% AI
Retired German Gradually classHuman (missed AI text)
Short control textGerman, Mixed
Sapling AI score100.0% AI
Retired German Gradually classHuman
Short control textEnglish, Human
Sapling AI score84.8% AI
Retired German Gradually classAI (false positive)
Short control textEnglish, AI
Sapling AI score100.0% AI
Retired German Gradually classAI
Short control textEnglish, Mixed
Sapling AI score86.6% AI
Retired German Gradually classAI

When reopened for this screenshot, Sapling displayed a prefilled provider example and an error. The screenshot is not an additional control and does not change the six documented results:

Sapling interface with a prefilled provider example and a visible error on September 20, 2026

The bars show Sapling’s own scores from 0 to 100. The labels state the documented source, so high scores for human controls are visible immediately.

German, Human
97.7%
German, AI
72.2%
German, Mixed
100.0%
English, Human
84.8%
English, AI
100.0%
English, Mixed
86.6%
Own short-text controls, September 19, 2026
gradually.ai

Both historical human controls received high AI scores from Sapling. The free check is therefore not enough on its own. A rounded score of 100.0% does not mean absolute certainty.

There is also a language limit. The official API documentation lists English only. A web interface accepting German text does not establish guaranteed German-language support. History and full-text storage were on in the observed controls. External visibility of the offered share certificates was not tested.

3.6 Gradually does not receive a pass for carrying this site’s name

The Gradually detector belongs to this website. That needs to be clear when its results are discussed.

This entry covers the German detector that has since been removed. It does not assess the English AI text detector that remains available on this site.

The German interface required no login and processed every short control. Its classifications still showed clear errors. In the full factual-text series, it labeled both human references as AI. In the separate short series, it made the opposite error too. The German AI text received the main Human class.

The reported confidence was usually 10%. The code defined that value as a floor. It is neither a measured certainty nor an AI share. A prominent result heading can look more decisive than the number beneath it supports.

This is what the German input page looked like before removal. The screenshot from September 20, 2026 does not show a test finding:

German Gradually detector interface with an empty text field before its removal

Result assignment also had usability problems. Replacing or clearing the text left the previous classification visible. History entries had disabled restore buttons. A reader could therefore connect an old result to newly pasted text.

I can’t recommend the retired German tool as a reliable checker on the basis of these results. The finding describes observed failures. It does not measure how often they occur overall.

4. Other services and research demos

A familiar product name does not always lead to a usable detector.

These ten entries combine provider claims with the result of access checks. If no text test ran, that is stated plainly. Images marked archival show an earlier interface. Their capture dates are not documented and their visible values do not demonstrate current detection performance.

4.1 Winston AI has a trial, but no measured result here

Winston AI lists German among its multilingual detection languages. The tested account showed 2,000 credits for 14 days, without a payment method.

That confirms access, not detection performance. No text was analyzed. The advertised document checker also receives no practical rating here.

The public Winston interface shows an empty input field. Its adjacent accuracy statement is provider marketing, not a measured result:

Winston AI product page with an empty detector field and provider marketing

The Winston pricing page also lists paid allowances. If you want to try it, start with known human writing. I do not have a measured detection result for Winston yet.

4.2 AI Detector Pro advertises German and three free scans

AI Detector Pro combines detection with rewriting functions.

According to its product page, the service supports German. Its free entry plan includes three scans per month. Humanization is outside the free allowance.

An existing login was rejected. I could therefore not analyze text, inspect reports or deletion functions, or confirm a paid-plan price.

The screenshot shows the public product page and a report illustration supplied by the provider. It is not from one of my scans:

AI Detector Pro product page with a report illustration supplied by the provider

If the same provider rewrites your text and then evaluates it, look closely at the result. That process does not replace an independent quality review.

See AI Detector Pro

4.3 Crossplag is now part of Inspera

Inspera acquired Crossplag and integrated its technology into its assessment offering. The official acquisition announcement also mentions AI text detection.

The former individual page led to a missing Inspera page in the final link check. A working route to the former standalone detector was not confirmed. Inspera’s current product was not tested.

The archival screenshot shows the former Crossplag interface, not today’s Inspera product:

Archival screenshot of the former Crossplag AI Content Detector with text input and result area

Claims about German plagiarism checks cannot be transferred to AI detection. I can’t rate Crossplag’s current AI detection from this evidence.

4.4 Kazan SEO now presents itself as Chadlinks

The Kazan domain now shows Chadlinks by Kazan SEO.

AI detection was still mentioned as an account feature during the review, but the existing login did not work. A current model, usable detector field, and reliable plan limits were not confirmed.

This is how the former AI GPT3 Detector looked. The archival image does not establish current access:

Archival screenshot of the former Kazan SEO detector with text input and a Real-Fake display

I would not treat that historical detector as a tested tool for current text.

4.5 Content at Scale now points to a Moxby writing-pattern helper

The former Content at Scale detector route redirects to a Moxby Marketplace mod.

It describes its result as an approximate writing-pattern aid. It is not intended to prove human or artificial authorship.

That limitation is sensible. It does not make the mod an unchanged successor to the former detector. I neither installed the required extension nor submitted text. Old Content at Scale results therefore do not belong under the current product name.

The archival screenshot explicitly shows the old detector, not the Moxby mod:

Archival screenshot of the former Content at Scale interface with a Human Content Score

4.6 Writer redirects the former detector route to its homepage

The former detector address redirects to the general Writer website.

No public input interface for the original product was visible there. The former free detector could not be confirmed through that route.

The archival image is the only illustration of the old checker. Its visible marketing claims are not verified current facts:

Archival screenshot of the former Writer AI Content Detector, not a current product interface

It remains unclear whether Writer has ended every detector function. Current enterprise products also do not establish that the old free checker is still available.

4.7 It is unclear who operates Unfluff today

The current Unfluff site describes a revival by enthusiasts or users. Its domain terms say the same.

An old WordPress plugin listing does not establish that the original backend still runs or who now operates it.

The archival screenshot shows an Unfluff Score, not a measured AI-detection rate:

Archival screenshot of Unfluff with text input and an Unfluff Score

Until the operator and backend are clear, I would not upload client writing or install a plugin there. Current AI detection was not tested.

4.8 Grover studies generated news text

Grover began as a research project for generating and detecting artificial news articles. Its official project page and source code remain available.

The linked hosted demo was unavailable when checked. That does not prove that the project has been permanently discontinued.

The archival image shows the former Detect tab, not a current test run:

Archival screenshot of the Grover research demo with an empty Detect input field

Grover is not currently tested as a browser service for day-to-day document checks.

4.9 The GPT-2 Output Detector is not a current ChatGPT test

The original OpenAI detector examined output from the GPT-2 era.

Its official Hugging Face Space was marked as paused. The original demo did not return a usable result. Available source code does not replace a current hosted-product test.

The archival screenshot illustrates the old GPT-2 demo. Its display does not belong to the current control series:

Archival screenshot of the GPT-2 Output Detector demo with a Real-Fake bar

A different detector on Hugging Face is not automatically the same model. Historical GPT-2 results also do not show how well a detector handles current AI writing.

4.10 GLTR shows word probabilities instead of a verdict

GLTR colors words according to how expected they are for a language model. The official repository explains the research approach.

A historical control with a German Kafka excerpt returned a histogram for 890 tokens. It used GPT-2 small. There was no native Human or AI class.

The archival screenshot uses a different English example. It does not document the Kafka control:

Archival screenshot of GLTR with three histograms and English text colored by word probability

It would be wrong to turn that output into an accuracy rate. The historical result also does not establish how well GLTR detects current AI writing. An insecure HTTP demo is not an appropriate place for confidential manuscripts.

5. Free and paid AI detector prices

The word limit is often more useful than the lowest monthly price.

These figures come from official product and pricing pages. They describe plan terms, not the scope of the tests. Dollar amounts are in US dollars. The Winston coupon prices are promotional, not list prices. Taxes and billing cadence can change what you actually pay.

Sources: official product and pricing pages, checked September 22, 2026. Free limits were not exhausted in every product.

ServicePangram Individual
Free allowance stated by provider2,000 words per day
Monthly billing$20 for 300,000 words per month
Annual billing$180 per year ($15 per month)
ServiceOriginality.ai Pro
Free allowance stated by provider3 scans per day, up to 2,000 words per scan
Monthly billing$14.95
Annual billing$155.40 per year ($12.95 per month)
ServiceCopyleaks Personal
Free allowance stated by providerFree allowance not confirmed
Monthly billing$16.99 for 100 credits: up to 25,000 words or 100 images
Annual billing$167.88 per year ($13.99 per month) for 1,200 credits: up to 300,000 words or 1,200 images
ServiceSapling Pro
Free allowance stated by providerWeb check up to 2,000 characters
Monthly billing$25
Annual billing$12 per month with annual billing
ServiceWinston Essential
Free allowance stated by provider2,000 trial credits for 14 days
Monthly billingRegular $18; $9 with displayed BTS50 coupon
Annual billingRegular $120; $60 with displayed BTS50 coupon
ServiceAI Detector Pro
Free allowance stated by provider3 scans per month
Monthly billingPrice not confirmed
Annual billingPrice not confirmed
ServiceGPTZero
Free allowance stated by provider10,000 words per month after sign-up, according to GPTZero
Monthly billingCurrent price not confirmed
Annual billingCurrent price not confirmed

With Sapling, keep its web plans and API apart. The API bills by character and has separate free limits. Sapling lists up to 100,000 characters per check for Pro web users. My web controls remained at 2,000 characters despite the visible Pro trial plan. The cause is unclear, so that observation does not establish a general paid-plan limit.

Copyleaks pricing bills Personal plans in credits. One credit covers up to 250 words or one image. Its 100-credit monthly allowance is not a free allowance. It can cover up to 25,000 words or 100 images.

The retired German Gradually detector was free. A usability control reached its daily limit after ten successful analyses.

Only choose an annual subscription after a relevant test series of your own. A lower price is little comfort if you must explain every false alarm later.

6. How to read an AI score without making a false accusation

Start with a text whose human origin you know.

Match its length, language, and text type to your normal work. An older post you wrote yourself is more informative than a random sample sentence.

If the detector already raises an alert there, you have found a real weakness. A second high score on client writing then carries much less weight.

Confidence, text share, and accuracy are different measures

These terms answer different questions.

Confidence describes how sure a model says it is about one classification. A text share describes marked portions. Accuracy measures correct decisions across a defined test set.

A display of 97% is not 97% accuracy. Averaging two provider scores does not create a truer result.

A majority vote from several detectors is not proof either. Systems can react to similar writing features, so several similar false positives can still be false positives.

Check the text and the writing process

When a result looks unusual, inspect the specific passage first.

Are the sources sound? Does the argument make sense? Can the author explain how the text came together? Version history and drafts help answer those questions without turning a score into an accusation.

Editorial review also needs different criteria. Human writing can still be poorly researched. An AI draft can become a useful starting point after proper fact checking and editing.

Do not use confidential writing as a trial input

Use non-confidential example documents for your first check.

Before uploading client material, review the specific product’s privacy information. How does it store text? Can it use content for training? Which deletion options and share links exist? An API promise does not automatically apply to the web interface from the same provider.

A free account, a private report, and a deleted history entry can all mean different things. Removing a visible list does not prove that every server-side copy disappeared.

7. My recommendation for your next document

For a second opinion, I would begin with Pangram after these controls. It labeled the recent control texts correctly. I would still account for the misplaced boundary in the mixed text.

Originality.ai, Copyleaks, and the retired German Gradually detector have documented false positives. Sapling returned high AI scores for human sources. GPTZero’s older run looked more nuanced but could not be confirmed with the newer text. The other active products do not yet have a measured performance judgment here.

Before you assess your next suspicious document, run a known human control through the selected tool.

You’ll learn where the tool becomes unreliable before the result affects a real person.

Frequently Asked Questions About AI Text Detectors

FH

Finn Hillebrandt

AI Expert & Blogger

Finn Hillebrandt is the founder of Gradually AI, an SEO and AI expert. He helps online entrepreneurs simplify and automate their processes and marketing with AI. Finn shares his knowledge here on the blog in 50+ articles as well as through the AI Business Club.

Learn more about Finn and the team, follow Finn on LinkedIn, join his Facebook group for ChatGPT, OpenAI & AI Tools or do like 17,500+ others and subscribe to his AI Newsletter with tips, news and offers about AI tools and online business. Also visit his other blog, Blogmojo, which is about WordPress, blogging and SEO.