Research
The citation gap: 12.4 sources a page, or 0.15
Across 3,038 competitor pages, the single measurement that separated researched writing from manufactured writing was how often a page linked to something outside its own domain.
We went looking for a measurement that would separate the pages in our corpus that were researched from the pages that were produced. We tried several that did not work before finding one that did, and the one that did is almost too simple to write up.
It is the number of links per page to a domain other than the publisher’s own.
The spread
Per-page averages, by cluster:
| Cluster | Pages | External links per page |
|---|---|---|
| One competitor’s blog | 36 | 12.4 |
| Another’s blog | 198 | 9.8 |
| Another’s blog | 104 | 8.5 |
| A features section | 15 | 12.7 |
| A 100-page blog | 100 | 0.15 |
| A 99-page how-to section | 99 | 0.0 |
| A 117-page use-case section | 117 | 0.0 |
| A 33-page comparison section | 33 | 0.0 |
That is a two-order-of-magnitude range within one industry, between companies of comparable size, on page types that serve comparable purposes.
Why it works when text statistics fail
Every other measurement we tried is computed from the words: length, unique words, shared passages, heading structure. Each of those describes how the text is arranged.
Arrangement is the thing generation is good at. You can produce text that is long, lexically unique, and structurally varied, in seconds, with no underlying work. Our own duplicate check scores the market’s largest generated cluster at 95% unique.
Citations are different in kind. A link to an external source is a residue of someone having gone and looked at something. It does not prove the person read it carefully, and it can be faked by anyone willing to spend ten minutes. But it cannot be produced by rearranging words, which is exactly the property we needed.
The honest caveats
It is not a quality score. A page can cite twelve sources and be useless. A brilliant personal essay may cite nothing. We are not proposing this as a universal measure of good writing.
It is domain-specific. It works here because the page types in question are comparison, research and documentation pages, which make factual claims about products that exist. A claim about another company’s pricing either came from somewhere checkable or was invented.
It can be gamed cheaply. If this measure became a target, pages would sprout decorative links. That is the usual fate of proxy metrics and we have no defence against it except not being big enough for anyone to bother.
What we did with it
We made it a publication condition rather than a score.
Our content schema already attached provenance to every fact: a source id, a retrieval date, a verbatim quote, and a confidence level. That was built for a different reason, to stop us making comparative claims about named companies that we could not defend.
Turning it into a gate was a small change. A page does not publish unless its factual claims carry sources. Our blog posts have the same requirement, which is why every post on this site ends with a list of the sources it drew on, including this one.
The nice property is that it is not a bar you can clear by writing better. Better writing does not produce a citation. Going and reading something does.