Skip to main content

Boardroom Answers · AI & Data · AI, Data & Analytics

I could build most of this with a foundation model and good prompts in a quarter. What is scientifically defensible here that a generic LLM wrapper doesn’t have?

The question a Chief Data Scientist / VP of AI & Analytics asks.

The short answer

A wrapper copies our prose, not our verification: provenance-gated numbers, consensus with adjudication, evidence-weighted confidence — and a calibration ledger where our forecasts are graded against reality, which no generic model can ever backfill.

The full executive answer

The honest answer is that a wrapper reproduces our prose and none of our verification, and verification is the product. Take any single output and count the deterministic machinery around the model: numeric provenance validation that rejects untraceable figures; hard-blocking guardrails on internal consistency; citation validation against a curated 74-framework catalogue; cross-provider consensus with structural diffing and third-model adjudication; evidence-weighted confidence bands computed server-side from a corroboration map the model never touches; a regression harness where known-bad outputs must be caught on every build. Each piece is unglamorous code — together they are a governed-output pipeline that took the better part of a year of failure-driven iteration, and none of it comes with an API key.

But the deepest moat is the one I would point a scientist to: the calibration loop. Every trajectory forecast the platform makes is logged as a scoreable prediction with a band and a due date; when the due date passes and a newer real assessment exists, the prediction is graded — absolute error, within-band or not — and the aggregate becomes a public calibration record. Our advice is audited by reality, conservatively: never scored without a real actual. A generic LLM structurally cannot close that loop, because it has no persistent relationship with the organisation’s measured future. Over time that record is either our strongest asset or our public humiliation — we built it anyway, because a decision platform that will not grade its own predictions is asking for faith, not trust.

Pre-launch honesty: the ledger starts near-empty, so today the defensibility is the verification pipeline and the loop’s existence; in two years it is the accumulated record no fast-follower can backfill. And the boring moats compound underneath: per-tenant evidence graded context, feedback converting into eval gold cases, and cost telemetry per generation — data assets a wrapper never accumulates.

Grounded in: Calibration methodology (predicted-vs-actual scoring, Brier-style within-band grading); NIST AI RMF MEASURE 2.5 (validity demonstrated over time); SR 11-7 outcomes-analysis expectations.

Want this answered live, on your data?