Skip to content

Base population

The base population is our model of the public: a large set of digital twins built to match the real composition of a country, from which every population you define is drawn.

You never run a simulation against the base population directly. You run against a population you have defined, and that population is a slice of the base, reweighted and often supplemented with your own data. So the base is the raw material. Its coverage sets the outer limit of who you can ask about, and its quality sets the floor on every answer.

How it is built

Real records first. The backbone is anonymised individual records from large-scale public studies — community and census surveys, expenditure diaries, time-use studies, long-running attitude surveys. These are real people who answered real questionnaires, and each comes with the statistical weight that says how many people like them exist. See Source.

Combined across studies. No single study covers circumstances, spending, time use and attitudes at once. Records are matched across studies for people of the same profile, so a twin can hold more than any one survey asked. Attributes brought in this way are labelled combined on the twin.

Filled out to match reality. Real records do not land evenly across every combination of age, income, region and household type. Where a combination is real but thinly covered, the base includes people drawn from the measured distribution for that combination rather than leaving a hole. Those twins are labelled, and groundedness reports how much of any population is made of them.

Deepened by interviews. Interview evidence gives some twins something no survey record has: their own account, in their own words. Interviews cannot cover a national population, so they are not the backbone — they are the part of the base with the most depth per person, and the part validation leans on hardest.

Weighted to be right in aggregate. The point of a base population is not that each twin is a perfect individual. It is that the group, taken together, has the composition of the real public: the right proportion of renters, the right income spread, the right regional mix. Weights are what make a sample of a few thousand speak for a country.

Versions

The base population has versions, and each is frozen. When new study data is released or new interviews land, a new version is built; existing versions stay exactly as they were.

This matters more than it sounds. A population is pinned to a base version, and a simulation is pinned to a population version. A result you ran in March still means what it meant in March, and a comparison between two runs is not quietly contaminated by the base having moved underneath them.

What it covers, and what it does not

The base covers adults in the markets we have built it for, at a level of detail set by what the underlying studies asked. It is deep on circumstances and behaviour, reasonable on attitudes, and thin on anything a public study has never had a reason to measure.

That last point has a practical consequence. If you define a population by something no study measures — a behaviour specific to your category, a relationship with your brand — the base cannot supply it. When that happens the builder tells you, and offers you the choice: use a named stand-in, tell us the shape of that group yourself, or bring data that carries it. What it will not do is quietly estimate the thing that defines your population and hand you a confident answer about a group that does not exist. See Population.

How to think about it

The base population is a map, not a census. It is built from real people and weighted to look like the country, and it is best understood as the most defensible available answer to "what is the public like" rather than as a directory of individuals.

The two numbers that tell you how far to trust it are groundedness, which says how much real evidence stands behind the people, and alignment, which says how closely they have matched real human answers when tested. See Validation.