Why two detectors disagree on the same text
Disagreement is a reading about the passage, not proof that one tool is broken.
Run the same paragraph through two detectors and the numbers rarely match. People tend to treat the gap as a puzzle with one right answer: one tool is accurate and the other is not. That framing is almost always wrong. Two tools disagree because they are answering slightly different questions, on different scales, against different reference sets. Once you see which parts of the pipeline differ, the gap stops looking like a malfunction and starts looking like information.
What this page is for: reading a detection score in context. It is not offered as a way past a checker, and no figure or example here should be taken as a promise about what any detector will report; scores move, and services disagree.
Where the two readings come from
- They are not measuring the same thing. One tool may weigh sentence rhythm and function-word balance, another may look at how predictable each token is. Different features produce different numbers even on identical text.
- They are calibrated on different populations. A model fitted on student essays and a model fitted on news copy place the same passage at different points on their own scales.
- Their scales are not comparable. Two tools can both print a percentage and still mean different things by it, so the numbers cannot be subtracted from one another.
- Their thresholds sit in different places. What one tool calls a flag near its cut point, another may call unremarkable simply because its default is more cautious.
- Length and register move each reading differently. A feature that is noisy on short passages in one tool may be stable in another.
- Agreement is not confirmation. Two tools built on similar features will often agree, and they can be wrong together; agreement is evidence about the tools, not about the writer.
How to read a disagreement
- Fix the input. Paste the same passage, with the same punctuation and paragraph breaks, into both tools before comparing anything.
- Record the direction of each score, not just its value. Note what it is meant to increase with, and what it claims not to cover.
- Look at the spread rather than the winner. A small gap is a stable reading; a wide gap says the passage sits where the tools do not agree, which is a statement about the text.
- Change one thing at a time. Re-run after trimming length or removing quoted material, so you can see which feature moved the number.
- Compare like with like. Score a second passage by the same writer in both tools; the pattern across passages matters more than any single run.
- Write down what you would have done if the tools had agreed. If the answer changes the outcome, the disagreement is a reason to slow down, not to reach for the friendlier number.
Questions people ask
Doesn't one of them have to be right?
Not necessarily. They are answering different questions. Both can give a reasonable reading of the same passage on their own terms, which is exactly why the numbers do not match.
Which tool should I trust?
Neither, on its own. A single tool is one observation. The useful question is whether independent evidence points the same way, not which number is larger.
Why do they agree on some passages and not others?
Passages in the middle of the range tend to produce agreement; passages near either end, or written in an unusual register, tend to produce the widest gaps.
Can I average the two scores?
No. Averaging numbers that are not on the same scale produces a third number that means nothing. Treat them as separate observations instead.
Is a large disagreement a sign that the text was edited?
It is a sign that the passage sits where the features diverge. Heavy editing, translation and formulaic structure can all produce that, and so can ordinary variation. The score alone cannot tell you which.