Research

The citation gap: 12.4 sources a page, or 0.15

Across 3,038 competitor pages, the single measurement that separated researched writing from manufactured writing was how often a page linked to something outside its own domain.

We went looking for a measurement that would separate the pages in our corpus that were researched from the pages that were produced. We tried several that did not work before finding one that did, and the one that did is almost too simple to write up.

It is the number of links per page to a domain other than the publisher’s own.

The spread

Per-page averages, by cluster:

Cluster Pages External links per page
One competitor’s blog 36 12.4
Another’s blog 198 9.8
Another’s blog 104 8.5
A features section 15 12.7
A 100-page blog 100 0.15
A 99-page how-to section 99 0.0
A 117-page use-case section 117 0.0
A 33-page comparison section 33 0.0

That is a two-order-of-magnitude range within one industry, between companies of comparable size, on page types that serve comparable purposes.

Why it works when text statistics fail

Every other measurement we tried is computed from the words: length, unique words, shared passages, heading structure. Each of those describes how the text is arranged.

Arrangement is the thing generation is good at. You can produce text that is long, lexically unique, and structurally varied, in seconds, with no underlying work. Our own duplicate check scores the market’s largest generated cluster at 95% unique.

Citations are different in kind. A link to an external source is a residue of someone having gone and looked at something. It does not prove the person read it carefully, and it can be faked by anyone willing to spend ten minutes. But it cannot be produced by rearranging words, which is exactly the property we needed.

The honest caveats

It is not a quality score. A page can cite twelve sources and be useless. A brilliant personal essay may cite nothing. We are not proposing this as a universal measure of good writing.

It is domain-specific. It works here because the page types in question are comparison, research and documentation pages, which make factual claims about products that exist. A claim about another company’s pricing either came from somewhere checkable or was invented.

It can be gamed cheaply. If this measure became a target, pages would sprout decorative links. That is the usual fate of proxy metrics and we have no defence against it except not being big enough for anyone to bother.

What we did with it

We made it a publication condition rather than a score.

Our content schema already attached provenance to every fact: a source id, a retrieval date, a verbatim quote, and a confidence level. That was built for a different reason, to stop us making comparative claims about named companies that we could not defend.

Turning it into a gate was a small change. A page does not publish unless its factual claims carry sources. Our blog posts have the same requirement, which is why every post on this site ends with a list of the sources it drew on, including this one.

The nice property is that it is not a bar you can clear by writing better. Better writing does not produce a citation. Going and reading something does.

WE ONBOARD EVERY TEAM OURSELVES

See it on
your product.

We’ll walk through CueFox against your own software, not a canned demo. You get a direct line to the people building it, and what you tell us shapes what ships next.

A person reads every request and replies. No newsletter, no sequence, no sharing your address.