Boardroom Answers · AI & Data · AI, Data & Analytics
Your verdicts carry confidence bands. Confidence in what, statistically? Is that a real interval or a vibes number the model made up?
The question a Chief Data Scientist / VP of AI & Analytics asks.
The short answer
The band is model-proposed, code-disciplined: hard-blocked if inconsistent, widened or lifted by measured evidence, its method printed on its face — and eventually graded against reality for coverage.
The full executive answer
It starts as a model-proposed range and is then disciplined by code — and the discipline is the honest part. First, validity gates: a guardrail hard-blocks any output whose band is inverted or outside zero-to-a-hundred — an internally inconsistent band is rejected outright, with one corrective retry and then failure; it never reaches a screen. Second, and more importantly, the band is adjusted server-side by measured evidence: our Evidence Map counts how many assessment dimensions are corroborated by independent sources versus merely self-reported, and unresolved conflicts mechanically widen the band downward while broad corroboration lifts the floor — with the adjustment reason appended to the band’s stated method, so you can read exactly why it is as wide as it is. The headline grounding number — what share of the verdict rests on independent evidence — is computed by our code from the map; the model is never asked for it, so it can never be invented.
What I will not call it: a frequentist confidence interval with coverage guarantees. It is a structured, evidence-adjusted uncertainty band — closer to a well-run estimation practice than to inferential statistics — and every band carries its method description on its face. The scientifically interesting commitment is that the bands become empirically testable: the calibration loop grades forecasts against later real outcomes, including whether reality landed within the stated band. So over time "is the band honest" stops being a design argument and becomes a measurable coverage rate. Pre-launch, that evidence is still accruing — but a platform that ships uncertainty on every verdict and then grades its own coverage is holding itself to a standard most human advisers never accept.
Grounded in: NIST AI RMF MEASURE 2.9 (uncertainty communication); ISO/IEC 23894 (risk information quality); calibration/coverage evaluation methodology.
The natural next questions
Related governed answers
- I could build most of this with a foundation model and good prompts in a quarter. What is scientifically defensible here that a generic LLM wrapper doesn’t have?
- Why no fine-tuned or proprietary model? You are renting your core capability from Anthropic — strategically, isn’t that fatal?
- If a board acts on a hallucinated number from your platform and loses money, who is liable — and can you prove, after the fact, that the number was not invented?
Want this answered live, on your data?