AI Check Writer

Human draft, machine polish: what a score sees

There is a difference between who had the thought and who typed the final sentence.

Between a person writing something and a machine writing something there is a large middle, and it is where most real documents sit. Someone outlines, argues and decides; a model tidies the prose. A score responds to the tidying and says nothing about the decisions, which is why arguments about detection rarely end with a number.

What this page is for: reading a detection score in context. It is not offered as a way past a checker, and no figure or example here should be taken as a promise about what any detector will report; scores move, and services disagree.

What changes when a model edits

Questions that get you further than a score

  1. Ask what the writer did, in order. Draft, then model, then human edit reads differently from model, then human, and the score often shows which shape it was.
  2. Separate the claim layer from the sentence layer. Who decided what the document argues is a question you can answer; who typed the sentences is one the score answers loosely.
  3. Check whether the edits changed meaning. Assisted editing that preserves the argument is a different case from editing that supplies it.
  4. Compare against an earlier draft if one exists. Version history is stronger evidence than any single reading of the final text.
  5. Apply the policy you already have, and apply it to everyone. If the rule is about declared assistance, enforce the rule; if it is about machine authorship, say what evidence would establish that.
  6. Write the reason down. A high score is not a reason, and the next person to read the file should be able to see what you weighed.

Questions people ask

If a model rewrote the sentences, is it AI-written?

That depends on the rule you are applying, and reasonable rules differ. What is not ambiguous is that the surface changed and a detection score will respond to that change. Whether the document is then treated as assisted or as generated is a policy question.

Can a detector distinguish polished human text from generated text?

Not reliably, because it does not measure the process. It measures how uniform the prose is, and a thorough polish and a generation both make prose uniform.

Does that make detection useless?

It makes it useful for what it is: a screen that points at passages worth a closer look, and a poor basis for a finding about a person.

What should a policy say about assisted editing?

Whatever the institution can check and apply evenly. A rule about declared assistance is enforceable; a rule about how much help is too much, measured by a score, is not.

What evidence settles the question?

Drafts, version history and the writer's own account. The score belongs beside those, not in place of them.

Where to go next