How to read a detection score
A score is an estimate about a piece of text, not a finding about a person.
Most disagreement about detection starts one step too late. By the time two people are arguing about a verdict, both have already accepted a number that neither of them read carefully. This page takes the number apart: what it is computed from, what it can and cannot represent, and the order in which to read it.
What this page is for: reading a detection score in context. It is not offered as a way past a checker, and no figure or example here should be taken as a promise about what any detector will report; scores move, and services disagree.
What the number is built from
- It is a measurement of text, not of authorship. The input is the passage in front of the tool; nothing about the writer, their first language, or how the passage was produced reaches the model.
- It is a position on a scale, not a probability. A percentage shows where a passage sits relative to a reference population, so it cannot be read as the chance that a machine wrote it.
- Both error types move together. Any setting that catches more machine text also catches more human text, so a score at the strict end is a statement about the setting, not about the passage.
- Length changes the reading. A few hundred words carry far less signal than a few thousand, and short passages swing for reasons that have nothing to do with the writer.
- The reference population is invisible. A tool calibrated on student essays reads like a tool calibrated on student essays, whether or not the passage in front of it is one.
- A number without a stated cut point means nothing. A value is not high or low until someone says which threshold they are applying, and why that one.
A reading order
- Read the direction before the value. Ask what the score increases with, and what it does not cover, before you look at the number.
- Ask for the unit of measurement. Was the whole document scored, or segments? A document figure is usually an aggregate that hides where the signal sits.
- Check length and register. Short passages, quoted material and formulaic text all behave differently, and none of them reflect on the writer.
- Run a second passage by the same author. Variation between passages is the fastest way to see how much of the score is about the writer and how much is about the text.
- Decide in advance what would count as evidence. If a score alone would not be enough to fail a person, it is not enough after the fact either.
- Write down the decision and the reason. The record is what makes the outcome reviewable by someone who was not in the room.
Questions people ask
Does a high score mean the text was written by an AI?
No. It means the tool placed the passage near the end of its scale that it associates with generated text. Real human writing lands there too, and how often depends on the language, the genre and how strict the threshold is set.
Can I read the percentage as a probability?
Only if you know how it was calibrated and on which population. In ordinary use it is a position on a scale, not the chance that something happened. Treating it as a probability is the most common misreading of a detection score.
Why does the same text score differently on two tools?
Because they measure different features and calibrate against different reference sets. Two tools agreeing is informative; two tools disagreeing tells you the passage sits in a region of the scale where the reading is unstable.
What threshold should I use?
One you can defend. A stricter threshold means fewer misses and more people wrongly flagged. Pick the side you can live with, state it before you look at the text, and keep it the same for everyone.
Should a score decide anything on its own?
No. It is one observation about a text. Decisions about people need a reason that a score cannot supply.