Skip to content

Identity

An identity is everything we know about one person, and where each piece of it came from. It is what turns a digital twin from a generic respondent into someone in particular.

If you have worked with marketing personas, this is the opposite of that. A persona is an archetype: a composite invented in a workshop, given a name and a photograph, standing in for a segment. An identity is a record about a real person, assembled from sources we can point at. Nothing in it is written to sound convincing.

What an identity contains

Four kinds of thing, and they carry different weight:

Circumstances. Age, where they live, household, work, income band, housing. The stable facts that place someone.

Behaviour. What they actually do. How often they order delivery, what they spend on groceries in a month, how they get to work, what they cancelled last year. Behaviour is the most useful part of an identity by a distance, because it is what a decision is actually made out of.

Stated views. What they have said about things — in an interview, a survey response, or their own public writing.

Their own words. Where we have them, verbatim. A line someone actually said does more work than three attributes describing them.

An identity holds a person's circumstances and behaviour, not a verdict about them. "Orders delivery two or three times a month, almost always with a promo code" is evidence. "Price sensitive: high" is an interpretation, and we leave those to the twin rather than baking them in. The difference matters because an interpretation written into an identity quietly decides the answer before the question is asked.

Observed, reported, estimated

Every line of an identity is labelled with how we know it:

LabelMeaning
ObservedIt appears directly in a source. A purchase, a survey answer, a recorded behaviour.
ReportedThe person said it themselves.
CombinedBrought together from more than one study of the same kind of person. Reliable at the level of the pattern; see Source.
EstimatedAssigned statistically because it was missing.

You can see these labels on any twin. They matter because they are the difference between "this person spends $180 a month on food away from home" and "people like this person typically do". Groundedness is the summary of that mix across a whole population.

Estimated values are used to make a population's composition correct — to get the right proportion of renters, or the right income spread. By default they are not part of what a twin is told about itself. A twin should not be told it earns a particular amount because its postcode suggests so.

What an identity is not

It is not a real person's file. Identities carry no names, no contact details and no addresses. A twin is a person-shaped set of facts, not a record you could use to find someone.

It does not change with the question. The same identity is presented to a twin whether you are asking about price, packaging or loyalty. This is a deliberate constraint: selecting the price-related evidence for a price question would produce a population that looks price-driven whenever you ask about price, which would be a reproducible finding and entirely an artefact of how we chose what to show.

It is not complete, and does not pretend to be. No record covers a whole person. Twins are told to fill in the gaps the way that person plausibly would, and never to contradict what is known about them. Where a twin has extrapolated, it shows up in its stated reasons, where you can see it and judge it — rather than being written into the identity, where you could not.

Reading one

Open any twin from a result and you see its identity in full, every line labelled with how we know it and which source it came from, alongside the decision it made and its reasons.

Reading ten of them after building a population is the fastest check available. If they do not read like your customers, the population is wrong in a way no summary statistic will tell you.