The four ways it goes wrong
Accuracy failures in this category are not random. They fall into four recognisable shapes, and each has a different tell.
- Invention. A language model asked for a figure it has no source for will supply one anyway, often to a convincing number of decimal places. The tell is a number with no dataset named beside it.
- Staleness. An answer drawn from training rather than from a document read today describes a world that may be two years old. The tell is the absence of a date on the finding.
- Echo. Ten outlets carrying one press release look like ten sources. The tell is a confidence claim with no independent-source count, or a count that does not fall when you ask which outlets were reprints.
- Silent omission. A tool that drops inconvenient material without saying so produces a tidy answer and an unauditable one. The tell is the absence of any record of what was excluded.
What to demand before you rely on an output
Five requirements, all of which a serious system can meet today. None of them is a courtesy; each one is what makes a claim checkable by somebody who was not there.
- A dated document behind every claim, openable in one click.
- A named dataset and period behind every number.
- A count of the independent sources carrying each finding, with reprints already merged.
- A record of what was held back, and why.
- A written statement of what the system does not do.
How to audit an output in ten minutes
Take any AI research output and do three things in order. The whole audit takes less time than reading the document properly.
First, pick three citations at random, open them, and check that each says what the output claims it says. Second, find the most surprising claim in the document and look for its corroboration count - a surprising claim with one source is a lead, not a finding. Third, look for the limits section. A research document with no stated limits has not been audited by its own author either.
Why "accurate" is the wrong word to argue about
Research is not a true-or-false exercise, and a system that claims otherwise is making a promise it cannot keep. What a research process can honestly offer is traceability: every claim leads to a document, every number to a dataset, and every disagreement between sources is reported as a disagreement rather than resolved behind the scenes.
That is a lower claim than accuracy and a far more useful one, because it is checkable. A reader who disagrees with a conclusion can open the same evidence and reach their own. A reader given an accuracy guarantee can only take it on faith.