The three layers of a research task
Most research work separates into three layers, and they automate very differently.
- Collection. Finding the sources and pulling the relevant figures out of them. Highly automatable, and the layer where most of the hours go.
- Synthesis. Turning a pile of material into a structured picture with the contradictions visible. Partly automatable, and the layer where automated output most often goes wrong quietly.
- Judgement. Deciding what the picture means for your company, and what to do next week. Not automatable, and the layer the automation exists to make room for.
Where automated market sizing is trustworthy, and where it is not
A market size can be built two ways. The defensible way is to name a specific dataset - a statistical agency series, a census, a national accounts figure - and state the period it covers, so a reader can open the same table and get the same number. The indefensible way is a confident figure with no series behind it, which is what a language model will produce if asked for a market size and given nothing to read.
This is the single most important thing to check in any automated research output. If a figure does not name its dataset and period, treat it as unsourced regardless of how precise it looks. Precision is not evidence; a number given to three decimal places is exactly as unsourced as a round one.
What automation is genuinely good at
Monitoring is where software wins outright. A person cannot check forty pricing pages, six job boards and a regulatory register every morning, and will not do it consistently for six months. Software will, and it does not get bored in week three, which is when manual monitoring quietly stops.
Citation is the second. Attaching the source to the claim at the moment the claim is made is mechanical work that humans do badly under time pressure, and it is the thing that makes a report survive being challenged.
What to ask a vendor
Three questions, and the answers are usually revealing: where does each figure come from, what did the system decide not to show me, and what does it do when the sources disagree. A system with good answers to all three is doing research. A system that cannot answer the second is doing summarisation.