AI Check Writer

How to read a detection score

A score is an estimate about a piece of text, not a finding about a person.

Most disagreement about detection starts one step too late. By the time two people are arguing about a verdict, both have already accepted a number that neither of them read carefully. This page takes the number apart: what it is computed from, what it can and cannot represent, and the order in which to read it.

What this page is for: reading a detection score in context. It is not offered as a way past a checker, and no figure or example here should be taken as a promise about what any detector will report; scores move, and services disagree.

What the number is built from

A reading order

  1. Read the direction before the value. Ask what the score increases with, and what it does not cover, before you look at the number.
  2. Ask for the unit of measurement. Was the whole document scored, or segments? A document figure is usually an aggregate that hides where the signal sits.
  3. Check length and register. Short passages, quoted material and formulaic text all behave differently, and none of them reflect on the writer.
  4. Run a second passage by the same author. Variation between passages is the fastest way to see how much of the score is about the writer and how much is about the text.
  5. Decide in advance what would count as evidence. If a score alone would not be enough to fail a person, it is not enough after the fact either.
  6. Write down the decision and the reason. The record is what makes the outcome reviewable by someone who was not in the room.

Questions people ask

Does a high score mean the text was written by an AI?

No. It means the tool placed the passage near the end of its scale that it associates with generated text. Real human writing lands there too, and how often depends on the language, the genre and how strict the threshold is set.

Can I read the percentage as a probability?

Only if you know how it was calibrated and on which population. In ordinary use it is a position on a scale, not the chance that something happened. Treating it as a probability is the most common misreading of a detection score.

Why does the same text score differently on two tools?

Because they measure different features and calibrate against different reference sets. Two tools agreeing is informative; two tools disagreeing tells you the passage sits in a region of the scale where the reading is unstable.

What threshold should I use?

One you can defend. A stricter threshold means fewer misses and more people wrongly flagged. Pick the side you can live with, state it before you look at the text, and keep it the same for everyone.

Should a score decide anything on its own?

No. It is one observation about a text. Decisions about people need a reason that a score cannot supply.

Where to go next