Skip to content

Confidence

Confidence is the label on a result telling you how far to trust it. It is not the model's opinion of its own answer. It is derived from how the population has actually performed against real people, plus how cleanly this particular simulation separated its options.

Every result carries one. It is the first thing to look at before you act on a number and the last thing to quote without.

What it is built from

Three inputs, all measured:

How this population has performed. Alignment: how closely these twins have matched real human answers on questions they had not seen. This is measured continuously, not asserted. See Validation.

How much the sample can resolve. The size of the sample and the shares that came out of it determine the range around each option. Weighted samples resolve a little less than their raw count suggests, and that is accounted for.

Whether the options actually separated. The gap between the top two options relative to the uncertainty around them. A simulation that cannot tell its leading option from its second is not high confidence no matter how well the population has scored historically.

A population that has never been validated caps confidence, and the result says so rather than showing an unexplained ceiling.

The levels

High. Trust the direction, the ranking and the rough magnitude. Act on it.

Moderate. Trust the direction and the ranking. Treat gaps between close options as unresolved — if the decision turns on those few points, get more evidence before committing.

Low. Treat it as a hypothesis worth testing rather than an answer. Usually a sign that the population needs more grounding, the sample was too small for the margins involved, or the options were too close to separate.

At every level, the exact percentages are approximate. Confidence is a statement about direction, ranking and magnitude, never about the second decimal place.

Confidence is not the twins' certainty

Twins also report how sure they each felt, and you can see that on a result. It is interesting — a 63% share where most twins said "could go either way" is a different situation from one where most were sure — but it is self-reported, and people are famously poor judges of their own certainty.

Confidence is measured from performance against real humans. Twin certainty is a texture in the result. Do not substitute one for the other.

Raising it

Improve the population before anything else. Groundedness and alignment are the largest inputs, and a better-grounded population raises confidence on every simulation you run against it, not just this one.

Increase the sample, if the problem is that two close options did not separate. This narrows ranges with diminishing returns, and it does nothing at all if the real problem is the population.

Fix the setup. A missing option in the action space or a scenario that is a topic rather than a situation will hold confidence down, and no amount of sample will fix either.

Run a comparison. Differences are on firmer ground than absolute levels, because both sides carry the same uncertainties. If a single result is moderate but the direction of a change is clear, the comparison may be what you can actually act on.