The short version: we ask twins questions that real people have already answered, compare the two distributions, and publish the result. The twins never see those questions or those answers while they are being built.
Held-out answers
Two kinds of real answers are held back and used only for testing.
From interviews. A portion of every interview round is withheld. Those people answer real questions, the population answers the same ones, and the two are compared.
From public studies. Long-running public surveys have asked thousands of questions of tens of thousands of people, with the results published. Those are a strong test set for two reasons: the questions were written by survey methodologists rather than by us, and there are enough of them to test across topics rather than on a handful we chose.
In both cases the held-back answers are excluded from everything used to build twins. A test the twins had already seen would measure nothing, so the exclusion is enforced at the point data is loaded rather than trusted to good intentions.
What is compared
The whole distribution, not just the winner.
A population that picks the right leading option for the wrong reasons scores badly, and it should: getting the shape right across every option is the thing that predicts whether it will hold up on a question nobody has tested yet. A population that gets "63/37" when the real answer was "63/37" scores well; one that gets "80/20" scores poorly even though it named the same winner.
The result is alignment, reported per population and refreshed continuously.
Why alignment is never one hundred
Ask the same person the same question a fortnight apart and they will not perfectly agree with themselves. That disagreement rate is the ceiling for anything trying to reproduce human answers, and a population claiming to beat it would be reporting a bug rather than a result.
Where we know the ceiling for a set of questions, alignment is shown against it, because "0.91 against a 0.94 ceiling" is a more useful number than "91%".
Drift
Populations are re-tested regularly rather than once. Alignment moves as the world moves, as new evidence arrives, and as populations are rebuilt. Drift is the change since the last test, and a population drifting downward is a signal to look at its sources before trusting its next result.
Alignment is not the whole picture
Alignment tells you how these twins have performed. It does not tell you how much real evidence stands behind them — a thinly built population can score well by luck on the questions tested and have no reason to hold on the next one. Groundedness is the other half, and both are shown on every population.
Against real behaviour
Held-out answers are one kind of evidence. What your customers actually did is better.
Where you connect outcome data, a past simulation can be scored against what really happened for the matching group over the matching period. That is a different and harder test than matching survey answers, and the scores are lower for good reasons — a simulation is a well-grounded estimate of how a group would behave, not a measurement of how it did. It is also the most defensible evidence available that any of this works, which is why it is worth setting up early.