Your content audit flagged eighteen things. Three mattered.
RankEcho's extractability checks flagged eighteen problems on a sixty-six block page. Four were real. The rest repeated one low-severity advisory or misread a deliberate editorial style as a defect. A checklist that produces eighteen findings when three matter teaches a writer to close the panel, so we cut the count by 78 percent without removing a single rule.
What did the checks actually flag?
A real page: a sixty-six block explainer, run through six deterministic extractability rules. The output was eighteen findings.
Reading them, four described genuine problems - three blocks too long to quote whole or carrying several numeric claims at once, and one advisory about figures buried in prose. The other fourteen were the same low-severity advisory repeated once per affected block, plus five instances of one rule misreading the page's style.
The four mattered. The fourteen were the reason a writer would stop reading.
Why does a long list of findings make a tool less useful?
Because attention is the scarce resource, not information. A panel showing eighteen items asks the reader to triage before they can act, and triage is the work the tool was supposed to do. The predictable outcome is that everything gets the same weight, which means nothing does.
The failure is quiet, too. Nobody files a complaint about a checklist being too thorough. They open it twice, find it exhausting, and stop opening it - and from the tool's side that looks like a feature nobody wanted rather than a feature that wore people out.
So the count is not a cosmetic property. A rule that is right and repeated fourteen times does more damage than a rule that is wrong once.
Which rule was wrong, and why?
The one checking whether section headings are phrased as questions. It fired five times on the same page - at blocks 5, 14, 16, 19 and 26 - because that page uses numbered section headings in a deliberate academic register rather than question form.
That was not a defect. It was an editorial choice, and the rule had no way to know because it looked at each heading in isolation. What it missed is that the page's own title is a question, so every section beneath it inherits the question it answers.
The fix was to give the rule that context: when the H1 is question-shaped, sub-headings are not flagged for failing to repeat the form. Five false positives disappeared and the rule kept working on pages whose titles label a topic instead of asking something.
How do you cut the count without weakening the checks?
Three changes, none of which removed a rule.
Group by block rather than by finding. A block with three problems is one thing to fix, not three things to read. The panel now shows one row per block, worst first, with every reason on that row.
Roll up repetition. A low-severity advisory triggering on five blocks becomes one line naming all five, because the writer's decision is the same either way - either they act on that advisory or they do not, and repeating it five times does not change which.
Give a rule the context it needs before flagging. The heading fix above is the example: most false positives come from a rule seeing less of the document than the reader does.
Eighteen findings became four. No rule was deleted, no threshold was loosened, and the three real problems are still the first three things on the panel.
Was the claim checking any better?
It was worse, and by a wider margin. The first version of the claim ledger reported forty-three factual claims on that page with fourteen lacking a source. Of the eight unsourced claims it displayed, one was a genuine unsourced statistic. Roughly thirteen percent precision.
The causes were specific and unflattering. The rule treated any sentence containing most or only as a comparative claim, so ordinary prose - the bias correction most people never hear about, the estimate needs only honest answers - was counted as an assertion requiring a citation. Headings were ledgered as claims, though a heading asserts nothing. And bare decimals were missed entirely, because the number pattern required a unit - so a genuine conservatism factor of 0.70 went uncounted while the word most was flagged five times.
After fixing those three, the same page reported thirty-eight claims with seven unsourced. Half the previous count, and the remainder is defensible.
What does this say about content audit tools generally?
Treat a finding count as a claim about the tool, not about your page. A checker reporting eighteen problems on a competently written page is more likely to be reporting its own thresholds than a defect in the writing.
Two questions separate a useful check from a noisy one. Does it tell you which block, by position, rather than describing a category of problem? And does it say why it fired, in terms you can disagree with? A finding you cannot argue with is a finding you cannot evaluate.
And a rule firing repeatedly on one page deserves suspicion before compliance. Five identical flags on a page written deliberately is more likely a rule missing context than an author making the same mistake five times.
Why publish our own false positives?
Because the alternative is asking you to trust a number we have not examined. This page argues that a finding count reflects the tool as much as the page - which would be a hollow argument if we had not counted our own.
There is a self-interested reason too, and it is worth stating. A tool whose output people stop reading has no effect regardless of how correct it is. Cutting eighteen to four was not a concession; it was the difference between a panel that gets used and one that gets closed.
The numbers in this piece are from development on our own content, not from a customer's site. They are small - one page, sixty-six blocks - and they describe our checks rather than checking generally.
What we still get wrong
The claim ledger cannot tell whether a source supports the claim beside it. It reports whether a link is present in the same block. A claim linked to a page that does not support it counts as sourced and should not.
The extractability rules measure form, not substance. A block can be the right length, open with an answer, carry a table and still say nothing worth quoting. No deterministic check reaches that, and pretending otherwise would be the same error at a different level.
And the heading fix has a known hole in the other direction: a page whose title is a question but whose sections genuinely wander now goes unflagged. That trade was deliberate - a missed finding costs less than a rule people learn to ignore - and it is a trade rather than a solution.
Frequently asked questions
Few enough to act on in one sitting. There is no correct number, but a long list on a competently written page usually says more about the tool's thresholds than about the page - and it produces the same outcome as no audit, because the reader stops reading.
A rule seeing less of the document than the reader does. Our worst example was a heading check that flagged five sections for not being phrased as questions, on a page whose title was already a question those sections answered. Giving the rule that context removed all five.
No, when the decision is the same either way. If a writer will either act on an advisory or not, listing it five times does not change the choice and does crowd out findings that need individual attention. One line naming all five blocks carries the same information.
Ask whether it names the block by position rather than describing a category, and whether it states why it fired in terms you can disagree with. A finding you cannot argue with is one you cannot evaluate, so it has to be taken on trust or ignored - and it will be ignored.
No. They come from development on one sixty-six block page of our own content and describe our checks at a point in time. The transferable part is the method - count your findings, read them, and treat a high count as a hypothesis about the tool.
