Write to us

AI Solutions · Service

Evaluating an AI system before you deploy it

The golden dataset, the harness, and what the score has to mean.

At a glance

Activity AI Solutions
Engagement Scoped mandate
Markets UAE · India · UK · US · Europe
Replies Two working days

What it is

“A demonstration is a curated sample shown by someone who wants a yes.”

A principal, Kaelo AI Solutions

Most AI projects are assessed by demonstration, which is the one method guaranteed to produce a favourable result. A demonstration is a curated sample shown by someone who wants a yes. An evaluation is a fixed dataset, a scoring method agreed in advance, and a number that can go down as well as up. The second is unglamorous and it is the only one that predicts what happens in production.

What’s included

Four areas of scope
01

A golden dataset that reflects reality

Assembled from real inputs including the difficult and ambiguous ones, labelled by people who know the domain, and frozen. A dataset curated to make the system look good is worse than no dataset, because it manufactures confidence.

02

Scoring chosen before results are seen

Exact match, semantic similarity, human rating or LLM-as-judge — each is appropriate somewhere and misleading elsewhere. Choosing after seeing the output is how evaluations become justifications.

03

Hallucination measured, not asserted

For retrieval systems, the question is whether an answer is grounded in the retrieved material. That is measurable, and it should be a reported number rather than a reassurance in a proposal.

04

Regression testing as the system changes

Models are updated, prompts are edited, retrieval is tuned. Without a harness that re-runs, you find out about a regression from a user.

How the work runs

STEP 01

Scope the task

What the system must do, and what would count as failing.

STEP 02

Build the dataset

Real inputs, domain labelling, difficult cases included.

STEP 03

Agree scoring

Method and threshold, written down before results exist.

STEP 04

Run, report, re-run

Including on every subsequent change.

When to come to us

  1. 01 You are being asked to approve an AI deployment on the basis of a demonstration.
  2. 02 A system is already live and nobody can say whether it is getting better or worse.
  3. 03 The output feeds a decision with a real cost attached to being wrong.

What we do not do

  • Evaluating on a dataset the builder assembled.
  • Reporting a single headline accuracy number without the difficult-case breakdown.
  • Publishing benchmark figures from someone else's system as though they were yours.

Common questions

Can we not just try it and see?
You can, and for low-stakes internal tooling that is often reasonable. Where the output feeds a decision that costs money to get wrong, 'try it and see' means discovering the failure rate through the failures.
Who labels the golden dataset?
People who know the domain — usually yours, not ours. Labelling is where domain knowledge enters the evaluation, and outsourcing it defeats the purpose.
How big does the dataset need to be?
Smaller than most people expect, provided it is representative and includes the hard cases. A few hundred well-chosen examples beat several thousand easy ones.
Begin

Send a brief. A principal reads it.

Written, considered replies within two working days.