AI Check Writer

What a detection score cannot see

The passage is the whole input, and everything outside it is invisible.

A detector is often described as if it were judging a writer. It is not. The entire input is the text in front of it, and everything outside that text is absent from the calculation: who wrote it, how, when, with what help, and under what constraints. Keeping that boundary in view is the difference between using a score as one observation and mistaking it for a judgement about a person.

What this page is for: reading a detection score in context. It is not offered as a way past a checker, and no figure or example here should be taken as a promise about what any detector will report; scores move, and services disagree.

The boundary of the input

What to supply yourself

  1. Write down what the score is about. One line: this is an observation about the passage, not about the writer.
  2. List what the tool could not see. Quoted material, translation, collaboration, revision history, constraints; whatever applies to this document.
  3. Supply the missing context yourself. Version history, notes, a second sample of the writer's work, or a direct question fills what the tool cannot.
  4. Separate observation from inference. Keep the score on one side and your reading of it on the other, and say which is which.
  5. Ask what would change your mind. If no further evidence would move the decision, the score has quietly become the verdict.
  6. Record both parts together. The observation and its limits belong in the same note, so a later reader sees the whole picture rather than just the number.

Questions people ask

So is the score useless?

No. It is one observation about the text, and observations are useful. It becomes misleading only when it is treated as a finding about a person.

Can a better tool see these things?

Not these. A tool that receives only a passage cannot see the writer or the process, however good its model is. That limit is structural rather than a matter of accuracy.

What about detectors that look at typing history?

Those are different instruments collecting a different input. They answer a process question rather than a text question, and they carry their own limits.

Why do translated passages often score high?

Translation tends to flatten idiom and produce predictable phrasing, which is the kind of pattern detectors associate with generated text. The person may still have written every word in their own language.

How should I describe a score to someone else?

As a reading of the passage, with its limits attached. Saying the tool placed this text near the machine end of its scale is accurate; saying the tool found that an AI wrote it is not.

Where to go next