Skip to main content

Boardroom Answers · AI & Data · AI, Data & Analytics

Your verdicts carry confidence bands. Confidence in what, statistically? Is that a real interval or a vibes number the model made up?

The question a Chief Data Scientist / VP of AI & Analytics asks.

The short answer

The band is model-proposed, code-disciplined: hard-blocked if inconsistent, widened or lifted by measured evidence, its method printed on its face — and eventually graded against reality for coverage.

The full executive answer

It starts as a model-proposed range and is then disciplined by code — and the discipline is the honest part. First, validity gates: a guardrail hard-blocks any output whose band is inverted or outside zero-to-a-hundred — an internally inconsistent band is rejected outright, with one corrective retry and then failure; it never reaches a screen. Second, and more importantly, the band is adjusted server-side by measured evidence: our Evidence Map counts how many assessment dimensions are corroborated by independent sources versus merely self-reported, and unresolved conflicts mechanically widen the band downward while broad corroboration lifts the floor — with the adjustment reason appended to the band’s stated method, so you can read exactly why it is as wide as it is. The headline grounding number — what share of the verdict rests on independent evidence — is computed by our code from the map; the model is never asked for it, so it can never be invented.

What I will not call it: a frequentist confidence interval with coverage guarantees. It is a structured, evidence-adjusted uncertainty band — closer to a well-run estimation practice than to inferential statistics — and every band carries its method description on its face. The scientifically interesting commitment is that the bands become empirically testable: the calibration loop grades forecasts against later real outcomes, including whether reality landed within the stated band. So over time "is the band honest" stops being a design argument and becomes a measurable coverage rate. Pre-launch, that evidence is still accruing — but a platform that ships uncertainty on every verdict and then grades its own coverage is holding itself to a standard most human advisers never accept.

Grounded in: NIST AI RMF MEASURE 2.9 (uncertainty communication); ISO/IEC 23894 (risk information quality); calibration/coverage evaluation methodology.

Want this answered live, on your data?