Quality Assurance
How we measure and protect rewrite quality.
Rewrite quality is the core product. This page documents how we measure it, the review loop it sits inside, and the specific standards a rewrite has to clear before we consider it good.
The five quality dimensions
Every rewrite is evaluated on five dimensions. We do not treat any of them as optional. A rewrite that scores well on four and poorly on one is not a good rewrite — it is a rewrite with a failure mode that will eventually surface in a real reader's response.
- Readability. The piece is easier to read than the original at the same level of substance — sentences scan, the structure is visible, the cognitive cost has gone down.
- Tone. The piece sits where it should on the four tone axes (warmth, directness, confidence, formality) for its actual audience, not for an average reader.
- Rhythm. Sentence length varies. The piece does not read as machine-uniform. Read aloud, it sounds like a person thinking.
- Resonance. The piece earns the reader's attention. It does not waste the opening on context; it does not waste the ending on summary.
- Clarity. The point is recoverable on a first read. A careful reader could state the piece's central claim back to you in one sentence.
The review loop
Rewrites that flow through the product are sampled continuously. A subset is reviewed by our editorial team against the five dimensions above. The reviewers are working writers and editors, not annotators with a checklist — the bar is whether the rewrite is something they would ship under their own name.
When a rewrite falls short on a dimension, the underlying pattern is logged. Patterns that recur become product work — model tuning, prompt refinement, or guardrail adjustments. Patterns that occur in isolation become editorial notes shared with the writers building our reference content.
What we do not measure with
We do not measure quality with AI-detector scores. Detectors report a probabilistic judgement that is uncorrelated with the things readers actually respond to. We do not measure quality with surface-level readability indices alone — Flesch-Kincaid is a useful signal, not a target. We do not measure quality with word-count compression ratios. Shorter is often better; shorter is not always better.
The only quality measure we trust is whether a careful human reader, looking at the before and the after, would say the rewrite is the version they would prefer to send.
How to report a poor rewrite
Every rewrite has a feedback control. We read every flag. If you can paste the before and after into a message to our team with a note on what went wrong, that is the highest-signal feedback we receive — and it is the kind that most reliably changes how the engine behaves on the next pass.
Related: editorial standards · research methodology · AI usage policy.