Automate in this order
Each stage is useful on its own, and each one makes the next easier.
- Monitoring. Watching a fixed set of companies and sources on a schedule. Highest time saved, lowest risk of a wrong answer.
- Collection and extraction. Pulling the figures and passages that matter out of documents you have already found.
- Assembly. Turning the collected material into a structured document with citations in place.
- Interpretation. Leave this with a person. It is where the value is and where automated output fails least visibly.
The citation rule
Bind the claim to the source at collection time. Concretely: when a figure is extracted, the document it came from is recorded with it, and that pairing is carried through every later stage rather than reconstructed.
The alternative - generating prose and asking a model to add citations - produces references that look right and sometimes are not. Once that has happened in a document that circulated, everything else in it becomes suspect too.
Checks before anything circulates
Four, and they take minutes:
- Open three citations at random. Do the sources say what the document says they say?
- Does every figure name a dataset and a period?
- Is anything marked unavailable, or did every question get an answer? A research output with no gaps is usually a research output that filled them.
- Do the sources disagree anywhere, and if so, does the document say so?
What to do about gaps
A gap is a finding. Where no public series covers a market, the honest output says so and explains what would answer it - a survey, an interview, a paid dataset. A system that never returns a gap is not more capable; it is less honest.