Boardroom Answers · Security & Compliance · Security & Cyber Risk
Banks have model risk management — validation, monitoring, challenge. You're selling AI judgment to boards. Where's your model risk framework?
The question a Chief Risk Officer (CRO) asks.
The short answer
Two frontier models independently produce and cross-challenge board-tier analyses; every number must trace to a source or the output fails; and a CI evaluation harness regression-tests AI behaviour on every commit — model risk management, translated to LLMs.
The full executive answer
We've borrowed exactly that discipline, translated to LLMs. Independent challenge: our highest-stakes outputs — board-tier analyses — run dual-provider consensus, meaning Anthropic's and OpenAI's models independently produce and cross-examine the analysis, with adjudication of disagreements; a single model's hallucination has to survive a second model's challenge before a board sees it. Validation: every AI response is validated against a strict typed schema — structurally wrong output is rejected, not rendered — and our numeric-provenance system enforces that every figure in generated prose traces to a licensed source; an unsourced number fails the output. Framework citations are validated against a curated catalogue, so the AI can't invent an authority.
Ongoing monitoring is where most AI vendors have nothing, and we have a CI-embedded evaluation harness: golden fixtures and mutation tests run through every registered AI schema's contract on every single commit — with zero live AI calls, so it's deterministic — meaning a prompt change or model-behaviour drift that breaks expected output shapes fails the build before deployment, not after a customer meeting. Cost and usage budgets per organisation catch runaway behaviour; every generation is provenance-logged (model, version, inputs) so any output is reconstructable after the fact — the model-risk equivalent of an audit file.
Honest boundaries: we do not train or fine-tune our own models — we govern frontier models from two providers, which means classic quantitative validation (backtesting a credit model against outcomes) doesn't map one-to-one; our controls are contractual (provider terms), architectural (consensus, schemas, provenance), and evaluative (the CI harness). And there is no outcome-tracking evidence yet that our recommendations correlate with better board decisions — we're pre-launch, and that longitudinal claim will take customers and time. The framework is real; the actuarial validation of it is future.
Grounded in: NIST AI RMF 1.0 (MAP/MEASURE/MANAGE); SR 11-7 (Fed model risk guidance, by analogy); EU AI Act Art. 9 (risk management system); OWASP LLM Top 10 LLM09 (overreliance).
The natural next questions
Related governed answers
- Every credible company can name its own top risks. What are yours — the ones that would actually hurt us as your customer?
- You deploy one codebase to all tenants. A single poisoned dependency is a breach of every customer simultaneously. How do you defend the software supply chain?
- Forget external hackers — your biggest threat is your own people. What stops one of your engineers, or frankly you yourself, from reading my board's data?
Want this answered live, on your data?