Confidence on a run is not a model’s self-report. It is derived from how that population has performed against humans it had never seen.
Held-out validation
A slice of every panel is withheld. Those people answer real questions, the population answers the same ones, and the two distributions are compared. The comparison is on the whole shape, not just the leading option, so a population that picks the right winner for the wrong reasons still scores badly.
Why alignment is never one hundred
Ask the same person the same question twice a fortnight apart and they will not perfectly agree with themselves. Human test-retest sets the ceiling, and a population that claimed to beat it would be reporting a bug rather than a result.
How to use the number
- 01HighTrust the direction, the ranking and the rough magnitude. Act on it.
- 02ModerateTrust the direction and the ranking. Treat the gaps between close options as unresolved.
- 03Medium or lowTreat it as a hypothesis worth testing, not an answer. Usually a sign the population needs more grounding.