What a detection score cannot see
The passage is the whole input, and everything outside it is invisible.
A detector is often described as if it were judging a writer. It is not. The entire input is the text in front of it, and everything outside that text is absent from the calculation: who wrote it, how, when, with what help, and under what constraints. Keeping that boundary in view is the difference between using a score as one observation and mistaking it for a judgement about a person.
What this page is for: reading a detection score in context. It is not offered as a way past a checker, and no figure or example here should be taken as a promise about what any detector will report; scores move, and services disagree.
The boundary of the input
- It does not know who wrote the passage. The writer's first language, training and history never reach the model; only the words do.
- It does not know how the passage was produced. Drafting, revising, dictating, translating and editing can all yield similar surface text, and the score cannot separate them.
- It cannot see quoted material as quoted. A passage full of citations may read as borrowed or formulaic, and the tool has no way to know why it does.
- It does not see the task or the genre. Formulaic forms reward a particular register, and that register looks similar whether a person or a model produced it.
- It misses collaboration. Text drafted by one person and finished by another, or assembled from notes, carries no mark saying so.
- It cannot report its own uncertainty about the writer. Everything the score expresses is a statement about the passage, and the gap between that and a person is exactly what it cannot cross.
What to supply yourself
- Write down what the score is about. One line: this is an observation about the passage, not about the writer.
- List what the tool could not see. Quoted material, translation, collaboration, revision history, constraints; whatever applies to this document.
- Supply the missing context yourself. Version history, notes, a second sample of the writer's work, or a direct question fills what the tool cannot.
- Separate observation from inference. Keep the score on one side and your reading of it on the other, and say which is which.
- Ask what would change your mind. If no further evidence would move the decision, the score has quietly become the verdict.
- Record both parts together. The observation and its limits belong in the same note, so a later reader sees the whole picture rather than just the number.
Questions people ask
So is the score useless?
No. It is one observation about the text, and observations are useful. It becomes misleading only when it is treated as a finding about a person.
Can a better tool see these things?
Not these. A tool that receives only a passage cannot see the writer or the process, however good its model is. That limit is structural rather than a matter of accuracy.
What about detectors that look at typing history?
Those are different instruments collecting a different input. They answer a process question rather than a text question, and they carry their own limits.
Why do translated passages often score high?
Translation tends to flatten idiom and produce predictable phrasing, which is the kind of pattern detectors associate with generated text. The person may still have written every word in their own language.
How should I describe a score to someone else?
As a reading of the passage, with its limits attached. Saying the tool placed this text near the machine end of its scale is accurate; saying the tool found that an AI wrote it is not.