Why market size figures fail in the room
The failure is rarely that a figure is wrong. It is that nobody can say where it came from. A board member asks where the number came from, the honest answer is "the report said so", and the entire slide loses its authority - along with the three slides built on it.
This happens because the market sizes in circulation mostly originate in commissioned reports whose methodology is behind the paywall, and are then quoted onward by people who never saw it. By the fourth repetition the figure has no provenance at all.
The five checks
In order. Any figure that fails one of the first three should not be used; failures on four and five can be worked around if you say so out loud.
- The dataset. Which specific series produced this number - a statistical agency table, a census, a national accounts series? "Industry sources" is not a dataset.
- The period. Which year or quarter does it measure, and when was the series last updated? A 2023 figure presented as current is a different claim from the one being made.
- The definition. Does the market as the series defines it match the market you mean? Statistical classifications draw boundaries that rarely match commercial ones, and this is where most honest errors live.
- The arithmetic. If the figure is a build - segments summed, or a share applied to a total - can you reproduce it from the stated inputs? If the inputs are not stated, it is not a build, it is an estimate.
- Measurement or projection. Is this a number somebody counted, or a number somebody extended forward? Both are legitimate; presenting the second as the first is not.
The five checks, run on a real series
Worth doing once end to end, because the five checks look abstract until a classification code is in front of you. Take a question somebody is actually asked in a board pack: how large is the business software market in Germany?
Check one, the dataset. Eurostat issues structural business statistics: annual enterprise statistics by NACE Rev. 2 activity, giving turnover, enterprise counts and employment per activity code per country. That is a named table with an identifier, which is what check one is asking for. Check two, the period. Structural business statistics run some years behind the current date, and the table states its reference year and the date it was last revised. Both belong beside the figure.
Check three, the definition, is where the work is. NACE J62.01 is computer programming activities: enterprises whose principal activity is writing software to order. A company selling a subscription product may sit there, or in J58.29, software issuing, and a firm doing both is counted once under whichever activity is larger. So the German business software market and turnover under J62.01 in Germany are two different quantities, and the distance between them is a judgement you have to state rather than a number you can look up.
Check four, the arithmetic. If you add J62.01 and J58.29 you have made a build, and a build has to name its inputs: both codes, both reference years, and the fact that some enterprises straddle the two. Check five, measurement or projection. Structural business statistics measure a year that has closed. Nothing in the table is extended forward, which is exactly why it can be checked.
No figure appears in this section, on purpose. Series get revised, and a number copied into a guide is stale by the time somebody quotes it back at you. The identifier and the activity code are the durable part; the figure belongs to the table, and reading it there takes a minute.
How to write the figure down
A market size travels. It leaves the analysis, enters a slide, is lifted into a board pack, and is quoted in a fundraise eighteen months later by somebody who has never seen the table. Five fields, written once, travel with it and answer the question before it is asked.
- The figure, at the precision the series states and no further.
- The dataset identifier, written as the agency writes it, rather than a description of it.
- The reference period, and the date the table was last revised.
- The market definition in one clause, in the classification’s own words.
- The date you read it.
Who questions the number, and what each one asks
The same figure draws three different challenges, and knowing which one is coming decides how the number should be presented.
- Finance asks whether the market definition matches the revenue line it is being compared against. A market drawn on a wider boundary than your own revenue makes your share look smaller than it is, and a narrower one flatters it. The share is only meaningful when both sides use one boundary.
- The board asks who measured it. An agency and a table is an answer. A report title is not, and the follow-up question arrives within the minute.
- An investor asks which part is measured and which part is judgement. A labelled ceiling with a stated addressable share reads as discipline. A single unlabelled number invites them to find the seam themselves, and they will.
Where the checkable figures actually live
For most markets a defensible figure exists in a statistical agency series, free to read, and it is less convenient than a bought report because it does not match commercial market boundaries. That inconvenience is the price of provenance.
- Europe: Eurostat structural business statistics, by NACE activity code.
- United States: the Economic Census for structure and the Bureau of Labor Statistics for employment and wages.
- Elsewhere: World Bank national accounts, and the national statistical office where one issues comparable series.
- Regulated sectors: the sector regulator usually issues counts nobody else has.
What to do when no agency measures your market
This is common for a new or narrow category, and there are two honest responses and one dishonest one.
The honest responses are to give the nearest wider sector as a labelled ceiling - "the market we sit inside is X, measured by this series; our addressable share of it is a judgement, not a measurement" - or to build from the bottom up with every input named, so a reader can disagree with an input rather than with the total. The dishonest response is to invent a figure with the precision of a measured one, which is what happens by default when a figure is demanded and no series exists.
What AI research changes, and what it does not
What it changes is the cost of provenance. Finding the right series, reading it, and carrying the dataset identifier and period through to the output is exactly the kind of work software should do, and doing it by hand is why people buy reports instead.
What it does not change is the judgement. A machine can tell you that the series defines the market one way and you mean another; it cannot decide which boundary is the right one for your decision. That stays with the person presenting the slide.
It also changes the shelf life. A figure with its dataset and period attached can be re-read when the agency issues its next reference year, and the difference between the two readings is itself a finding. A figure whose provenance was a report title cannot be re-read at all, which is why those numbers survive in decks long after they stopped being true.