Skip to content

Case study · Enterprise

Raising AI accuracy for an enterprise running AI in production

Client: Multi-cloud enterprise (anonymised)

Challenge

Where they started.

AI assistants and agents deployed on GCP, AWS and Azure with no systematic evaluation. Accuracy drifted with every model or prompt change.

What the FDE team did

Inside their stack.

  • Built golden datasets from real traffic
  • Evaluation harness tracking F1, precision, recall, groundedness and hallucination rate
  • Benchmarked multiple models per task
  • Added regression gates to CI/CD
  • Set up continuous eval dashboards
  • Ran weekly improvement cycles

Architecture

How it fits together.

Production traffic
Golden datasets
Eval harness
CI/CD gate
Model router
F1PrecisionRecall· 12 weekly eval cycles (illustrative)
0.50.751

Model comparison · placeholder

ModelF1p95Cost/1K
Claude[0.XX][X]ms$[X]
GPT[0.XX][X]ms$[X]
Gemini[0.XX][X]ms$[X]
Llama (Groq)[0.XX][X]ms$[X]

Outcomes · illustrative placeholders

The scoreboard.

[0.XX]→[0.XX]

F1 score

Illustrative · placeholder

+[X] pts

Precision

Illustrative · placeholder

+[X] pts

Recall

Illustrative · placeholder

−[X]%

Inference cost via model routing

Illustrative · placeholder

Google CloudAWSAzureClaudeOpenAIGeminiLlamaGroq
[Client quote placeholder — awaiting approval]

Stop building technology. Start engineering results.

Start with an outcome audit. We’ll map the systems, measures and FDE pod needed to move from AI ambition to production impact.

Deploy an FDE team