Skip to main content

Boardroom Answers · Security & Compliance · Security & Cyber Risk

Banks have model risk management — validation, monitoring, challenge. You're selling AI judgment to boards. Where's your model risk framework?

The question a Chief Risk Officer (CRO) asks.

The short answer

Two frontier models independently produce and cross-challenge board-tier analyses; every number must trace to a source or the output fails; and a CI evaluation harness regression-tests AI behaviour on every commit — model risk management, translated to LLMs.

The full executive answer

We've borrowed exactly that discipline, translated to LLMs. Independent challenge: our highest-stakes outputs — board-tier analyses — run dual-provider consensus, meaning Anthropic's and OpenAI's models independently produce and cross-examine the analysis, with adjudication of disagreements; a single model's hallucination has to survive a second model's challenge before a board sees it. Validation: every AI response is validated against a strict typed schema — structurally wrong output is rejected, not rendered — and our numeric-provenance system enforces that every figure in generated prose traces to a licensed source; an unsourced number fails the output. Framework citations are validated against a curated catalogue, so the AI can't invent an authority.

Ongoing monitoring is where most AI vendors have nothing, and we have a CI-embedded evaluation harness: golden fixtures and mutation tests run through every registered AI schema's contract on every single commit — with zero live AI calls, so it's deterministic — meaning a prompt change or model-behaviour drift that breaks expected output shapes fails the build before deployment, not after a customer meeting. Cost and usage budgets per organisation catch runaway behaviour; every generation is provenance-logged (model, version, inputs) so any output is reconstructable after the fact — the model-risk equivalent of an audit file.

Honest boundaries: we do not train or fine-tune our own models — we govern frontier models from two providers, which means classic quantitative validation (backtesting a credit model against outcomes) doesn't map one-to-one; our controls are contractual (provider terms), architectural (consensus, schemas, provenance), and evaluative (the CI harness). And there is no outcome-tracking evidence yet that our recommendations correlate with better board decisions — we're pre-launch, and that longitudinal claim will take customers and time. The framework is real; the actuarial validation of it is future.

Grounded in: NIST AI RMF 1.0 (MAP/MEASURE/MANAGE); SR 11-7 (Fed model risk guidance, by analogy); EU AI Act Art. 9 (risk management system); OWASP LLM Top 10 LLM09 (overreliance).

Want this answered live, on your data?