glossary · ai content detection
What AI content detection measures, and what it cannot see
Detectors return a probability that a model wrote your text, inferred from the words alone. What they actually measure, why false positives are common, and why unoriginality rather than authorship costs you traffic.
AI content detection is the attempt to judge from text alone whether a language model wrote it, usually by scoring how statistically predictable the wording is, and always producing a probability rather than proof of authorship.
reviewed 2026-09-02 · by the IT Master editorial team · how we check facts
What AI content detection is in practice
A detector reads a passage and returns a probability that a language model produced it. Nothing in the text records who typed it, so the tool is inferring from the words alone.
Most of them work the same way. They measure how predictable the writing is to a language model of their own: how often the next word is the obvious one, and how little sentence length and structure vary. Human drafts tend to be lumpier — an odd word choice, a clause that runs long, a fact that arrives out of order. Model output sits closer to the average. The gap between the two is what becomes a percentage.
Two consequences follow from the method itself, not from any one product being poor.
- The output is a probability, not evidence. "98% AI" means the text looks like the smooth part of a distribution. Nobody observed a model writing it.
- The signal thins out under editing. Rewrite a few sentences, add a table of your own figures, and the score moves, whether or not a model wrote the first draft.
Watermarking is a separate thing: a provider marks its own output at generation time. That is a signal from the producer, not a verdict from a reader.
AI content detection is the attempt to judge from text alone whether a language model wrote it, usually by scoring how statistically predictable the wording is, and always producing a probability rather than proof of authorship.
Why it matters in 2026
Detectors do not set rankings. No search engine publishes a per-URL detector score, and none appears in Search Console. What a detector decides is whether a person lets your page through: an editor at a client, a marketplace reviewing a listing, a procurement team that pasted your case study into a checker.
So the risk is commercial rather than algorithmic, and it runs both ways.
- False positives are common enough to plan for. Plain, well-organised prose — exactly what a specification page or a fitting procedure demands — reads as predictable, which is the property these tools score as machine-like. Writers working in a second language get flagged for the same reason.
- A clean score proves nothing useful. Text can be shuffled until any detector passes it while it still paraphrases the pages that already rank.
The thing that actually costs traffic is unoriginality, not authorship. A page that repeats what is already ranking gives no search engine or assistant a reason to select it, whoever wrote it. That standard is what helpful content describes, and it is a far better target than a detector percentage.
How this engine handles it
AI-tell detection is one of four checks every draft faces before it can publish, alongside fact-check, novelty scoring against the competing pages, and an E-E-A-T critique. None of them is run by the model that wrote the draft: Claude models write, and the fact-check and novelty judges are Gemini and GPT, a different vendor. A failing draft comes back with specific fixes, up to three rounds. One that still fails is not published, and a draft that fails the checks is not charged at the full rate — billing is per published article, see pricing.
Be clear about what that check is. It hunts the tells — the throat-clearing opener, the three-part list in every other paragraph, a hedge where a number belongs — and sends them back to be removed. It is not a promise about what any third-party detector will report. Editing a page until a classifier is satisfied optimises for the classifier, not the reader.
The durable part is grounding. Articles are built on the customer's own first-party data — product records, supplier datasheets, the questions support actually receives — material no rewrite of a competitor's page can contain. Measured, not promised: across 186 Standard-tier runs, 51% of drafts passed every check first time; the rest were revised or rejected.
Common misunderstandings
"A high score is proof." It is a classifier's guess about predictability, and it names no source, no draft history and no author. That is why serious editorial policies ask for drafts, notes and version history instead of a percentage.
"A humaniser fixes the problem." Rewriting tools that roughen the rhythm and swap in unusual synonyms can move a score without adding a single fact. You end up with worse prose that still says what the competing pages said.
"Google runs a detector on my site." Google's published spam guidance addresses unoriginal content produced at scale to manipulate rankings — the output, not the tool that typed it. There is no authorship verdict to appeal.
"Detection and watermarking are the same." One infers from the text after the fact; the other is a mark a provider applies to its own output at generation time, readable only where that provider's scheme is supported.
"Passing means the page is good." A detector cannot tell whether your figures are right, your sources exist, or your page adds anything. That is the job of cross-model validation and, ultimately, of a named editor.
questions people ask
Are AI content detectors accurate?
They are estimates, and their errors run in both directions. Because they score how predictable the wording is, clear technical writing and prose from second-language authors are flagged as machine-written far more often than the marketing claims suggest, while lightly edited model output frequently passes. Accuracy also drops sharply on short passages, where there is too little text to judge. Treat any figure a detector reports as one weak signal among several, never as evidence of who wrote something, and never as grounds for an accusation on its own.
Does Google penalise AI-written content?
There is no per-URL detector score, and nothing in Search Console reports one. Google's published spam guidance is about unoriginal content produced at scale to manipulate rankings, which describes the output rather than the tool that produced it. In practice, pages fall away when they repeat what is already ranking and give a reader no reason to stay. The workable response is to add what only you hold — your own data, specifications, prices, support questions — and to have facts checked before publishing, rather than to chase a detector percentage.
Should I run articles through a detector before publishing?
As a rough smoke alarm, occasionally. As a gate, no. Editing a page until a classifier is satisfied optimises for the classifier, and the changes it rewards — more unusual words, choppier sentences — rarely make the page more useful. A better pre-publish routine asks different questions. Is every figure sourced and dated? Does the page contain at least one thing a competitor cannot copy? Did someone other than the writer check the facts? A page that clears those reads well to a person, whatever a detector says about it.
Can AI writing be made undetectable?
Chasing that goal is the wrong project. Scores can be pushed down by rewriting, adding original data and varying structure, but detectors change, and a passing score still says nothing about whether the article is worth reading. The property worth engineering is originality: information that exists only in your business, claims checked by something other than the model that wrote them, and an accountable editor on the page. That is also what search engines and assistants select for, so the effort pays twice.
See what search engines and AI assistants find on your site
Free, no account. Type your address and we show you what is missing and what we would write first.