Intelligence · Archive

The archive no one else can just copy.

A bigger dataset is easy to fake. A defensible one isn't. Here's what it actually took to make one colour archive that a competitor can't shortcut their way past — and why the same discipline applies whether the source is a 200-year-old book or an AI-generated paragraph from thirty seconds ago.

Anyone can ship a bigger colour database than us. Scrape a dataset, wrap an API around it, call it done — for a while, it'll even look like more value. It isn't, and it never survives contact with the one question that actually matters: where did this number come from, and could you prove it if someone checked? Most colour tools have no answer. We built the entire product around making sure we always do.

That's not a slogan. It's cost us weeks at a time, on multiple archives, because the alternative — the fast way, the way a competitor would take — turned out to be indefensible the moment we looked closely.

The data that didn't survive the check

Robert Ridgway's Color Standards and Color Nomenclature (1912) is a foundational reference in ornithology and natural history — 53 plates carrying more than a thousand mounted colour samples. We found a hex-value dataset for it already circulating. Rather than treat those values as source data, we went back to high-resolution scans of the 1912 plates and independently extracted the colours ourselves: colour-calibration targets, cross-scan corroboration, the works. Only then did we compare the two datasets.

The result was the thing that stopped us: the values weren't merely close. They were identical across the corpus. That raised an obvious provenance question — were these really independent measurements, or did both datasets share the same underlying derivation? The exact agreement was a provenance red flag, not proof of anything by itself, but it was reason enough not to have built on those numbers in the first place.

Ridgway himself, it turns out, went to remarkable lengths to make independent verification possible at all. His own preface explains that every one of the 1,115 colours was "painted uniformly on large sheets of paper from a single mixture of pigments, these sheets being then cut into small squares" — specifically so no two copies of his book would ever vary. When we compared our own extraction against the dataset we'd walked away from, there were zero exact hex matches. That wasn't proof of independence on its own — the independent workflow was — but it's consistent with what we'd expect from two genuinely separate digitisation efforts.

Werner & Syme's Nomenclature of Colours (1821) — the book Darwin consulted aboard the Beagle — turned out to be a more interesting case than "the other one wasn't standardised." Syme, it turns out, tried just as hard as Ridgway did, nearly a century earlier: each colour was painted onto a large sheet, cut into individual swatches, and pasted into every copy of the book — the same production logic, applied to try to guarantee the same result. Both men were making a serious attempt at a reproducible standard.

Two hundred years of material history defeated the standard anyway. Every surviving copy is now its own physical object, with its own history of light exposure, storage, handling, moisture and ageing. A colour that began life on the same painted sheet in 1821 does not necessarily arrive on a twenty-first-century scanner looking the same as its siblings. So instead of picking one scan and calling it truth, we read three independent copies — Getty, Smithsonian, McGill — and reported what they actually said, including where they disagreed. Across all 110 colours, that comparison didn't justify treating any single hex as authoritative. So that's what we publish: a real range, with the disagreement stated, instead of the false precision of a single authoritative hex.

We also tried to recover the book's original animal, vegetable and mineral reference notes — charming period detail like "the egg of a grey linnet." The geometry of extracting that text worked perfectly. The optical character recognition on the reference columns printed alongside each plate's swatches — a different, more cramped typesetting than the book's ordinary prose, and the hardest text in the book to read off a scan — did not. When we checked the words two scans supposedly agreed on, we found things like "Egg of Grey flannel" and "Eqg of Grey Linnet" — two documents making two different mistakes, occasionally lining up by accident. Two wrong answers that happen to rhyme are not corroboration. So we shipped the colours without that detail, and said plainly why, instead of quietly filling the gap with something that read fine but wasn't real.

Two wrong answers that happen to rhyme are not corroboration.

A standard is not the same thing as eternal truth. Both Ridgway and Syme made serious, careful attempts at reproducibility — and physical evidence still has a history after it leaves its maker: light, moisture, storage, handling, conservation and scanning all become part of the chain. Colour Memory's job isn't to pretend that chain doesn't exist. It's to preserve it, and say what it did to the evidence.

The same test, pointed at thirty seconds instead of two hundred years

The discipline doesn't stay in the archive. We hold AI-generated content — including our own drafts — to exactly the same standard, and it catches things just as often.

A marketing mockup once cited a specific government regulation: a real-sounding document number, a real-sounding rule number, precise and official-looking. Both were wrong. A second version of the same mockup, generated moments later, cited a different wrong regulation number for the same fact. Two authoritative-looking citations, both invented, disagreeing with each other. Another time, a proposed "historical colour pairing" quietly reused two hex codes from an unrelated marketing image generated earlier in the same conversation — nobody faked a data point on purpose, it just pattern-matched something plausible-sounding into the shape a real citation would take.

Confident and specific is not the same as true, whether the source is a two-century-old book or a paragraph generated thirty seconds ago. Colour Memory isn't built on the assumption that old sources are true and AI is unreliable. It's built on the assumption that every claim — old, new, human or machine-generated — has to earn the confidence attached to it.

Why this is the moat, not just the ethics

Here's the part that should worry anyone trying to compete with a spreadsheet and a nicer interface: the check that made us reject an inadequately sourced dataset works on us too. Corpus-wide identity against an independent re-derivation is a provenance signal either way. A competitor who scraped Colour Memory's output and repackaged it would leave exactly the kind of provenance signature we'd know to investigate.

Which means there's no cheap way in. Nobody can shortcut this by hiring a scraper. The only way to actually compete with a real, independently-sourced archive is to either license it properly or go and do the underlying work — three copies, calibration targets, honest refusal when the evidence doesn't hold. That's slow and expensive by design. It's what separates a real archive from a copy of one, and it's the reason a bigger competing database isn't actually a threat: size was never the hard part.

  • Documented — we can point to the actual source.
  • Estimated — real evidence, reconstructed rather than directly measured, and labelled as such.
  • Interpretive — a reasonable historical or cultural reading, clearly marked as one.

And when we don't know, we say that too — an honest gap, not a confident guess dressed up to look like one. Most colour tools will hand you a hex code and let you assume it means something. We'd rather tell you exactly how much it means, and let anyone who wants to check, check.

Select a colour. See what we can actually prove.

← Back to blog