Boardroom Answers · AI & Data · AI, Data & Analytics
Why no fine-tuned or proprietary model? You are renting your core capability from Anthropic — strategically, isn’t that fatal?
The question a Chief Data Scientist / VP of AI & Analytics asks.
The short answer
We rent the commodity — frontier reasoning — and own everything defensible around it: context, verification, evals, calibration. Two providers in production keep the exit door open, and our own harness is the test any replacement model must pass.
The full executive answer
It is a deliberate architecture bet, and I will defend it on the merits. Our differentiation is not next-token quality — it is the governed context that goes in and the verification pipeline the output must survive: evidence-graded organisational data, a 74-framework knowledge base, provenance and guardrail gates, consensus, calibration. Fine-tuning would buy us a static snapshot of a frontier model’s reasoning, an evaluation burden we would carry alone, a serving cost, and a model that goes stale every time the frontier moves — which has been every few months. Renting the frontier means every provider capability jump flows into our product at the cost of a configuration change; the model registry and complexity-tier routing were built so a swap forces a governance entry, not a rewrite.
The dependency risk is managed rather than denied: two independent providers wired in production — Claude primary, GPT-4o as automatic fallback and as the consensus counterweight — with circuit breakers, so no single vendor outage or deprecation strands us; and the deterministic analytics keep working with no model at all. Strategically, everything defensible we build is model-agnostic on purpose: the eval fixtures, the feedback-derived gold cases, the calibration record, the tenant evidence graph. If the calculus ever flips — provider economics, data-residency demands, or an open-weights model clearing our eval bar — those assets transfer intact to a self-hosted or fine-tuned model, and the harness is exactly the instrument that would prove the candidate is good enough. We keep the option; we refuse the premature commitment.
Grounded in: NIST AI RMF MAP 1.1 (context-appropriate design decisions, documented); ISO/IEC 42001 supplier-management controls for third-party AI; model cards discipline making swaps governable.
The natural next questions
Related governed answers
- I could build most of this with a foundation model and good prompts in a quarter. What is scientifically defensible here that a generic LLM wrapper doesn’t have?
- Your verdicts carry confidence bands. Confidence in what, statistically? Is that a real interval or a vibes number the model made up?
- You are concentrated on Anthropic. When Claude has an outage or deprecates your model mid-quarter, what happens to my board meeting?
Want this answered live, on your data?