How we found our own error: a margin that could not exist
The most dangerous kind of data error is the one that produces a plausible answer. Here is one we shipped, how it survived scrutiny, and what finally caught it.
The symptom that looked normal
For months the highest-operating-margin ranking on this site was led by property trusts and equipment-rental companies. Nobody flagged it, including us, because it is what a reader expects: those businesses really do run high margins, and a list full of them reads as a correct list. The ordering was not evidence of anything. It was the error describing itself.
The number that could not exist
What finally broke it was a different entry in the same list: a bank with an operating margin above 200%. Operating income is revenue minus the cost of earning it, so a margin above 100% is not a surprising result, it is an impossible one. Either the profit is too large or the revenue is too small, and profit figures rarely go wrong in that direction. That pointed at revenue.
Pulling the thread at the source
Every figure on this site comes from a company's own XBRL filing, so the check was to go back to those filings and read the raw tags. One equipment-rental company had filed two numbers for the same year: $3.7 billion under the tag introduced by the 2018 revenue-recognition standard, and $16.1 billion under the older, broader Revenues tag. Both were correct filings. The first one simply does not include rental income, because rent is not revenue from a contract with a customer under that standard. This site had been reading the first one.
That single preference had been quietly understating revenue for every insurer, bank, property trust and rental company in the catalogue — and because margin is profit divided by revenue, a revenue that is too small produces a margin that is too large. The REITs at the top of the list were not there because they are profitable. They were there because their revenue was missing.
What changed
The selection rule now reads every candidate tag and takes the largest, because these tags are nested rather than interchangeable: a subset can never be a company's total. Eighty-four figures moved; the biggest corrections were insurers, whose premiums had been left out entirely. The tag that won is printed under each figure, and the full list of changed figures is published with the data.
Three further rules came out of the same investigation. Margins above 100% are now excluded from rankings outright rather than merely capped at a generous level, because the old ceiling was hiding exactly this class of error. A company whose reported operating income is at least as large as its revenue is dropped from every revenue-based ranking, not just the one where the impossibility showed up — if the denominator is wrong, every ratio built on it is wrong. And gross profit is no longer computed as revenue minus cost of revenue, because for many filers the cost tag holds only part of the cost.
The general lesson
If the top of a ranking is clustered by industry, suspect the pipeline before you admire the companies. Real leaderboards are usually mixed; a homogeneous one often means a single rule is selecting for an artefact rather than for performance. It is a cheap check and it costs nothing to run on somebody else's data, including ours.
Every company page links to the filing behind each figure, and the machine-readable archive published with each data refresh lists every figure the tag change moved, with its old and new value. Corrections are logged rather than quietly overwritten.