Write to us

AI SolutionsService

Evaluating an AI system before you deploy it

The test dataset, the scoring method and what the score has to mean.

Kaelo Global is a Dubai company licensed in Meydan Free Zone, with over 100 clients served so far.

At a glance

ActivityAI Solutions
EngagementScoped mandate
MarketsUAE, India, UK, US and Europe
Reply timeWithin two working days

What it is

A demonstration is a chosen sample shown by someone who wants a yes.

A principal, Kaelo Global AI solutions

Most AI projects are judged by demonstration, which is the one method almost guaranteed to produce a favourable result. A demonstration is a chosen sample shown by someone who wants a yes. An evaluation is a fixed dataset, a scoring method agreed in advance and a number that can go down as well as up. The second is less exciting, and it is the one that predicts what will happen in production.

A useful AI evaluation harness has three parts: a dataset built from real inputs including the difficult ones, a scoring method chosen before anyone sees the results, and a way to run the same test again whenever the model, the prompt or the retrieval changes. Without the third part, you learn about problems from your users.

We build evaluations for systems that feed real decisions, such as document processing, customer replies and internal search, and we are happy to test a system that someone else has built.

Tell us about the process.

It helps us understand your business before we reply.
We use it only to reply to this enquiry, usually on WhatsApp.
We send a copy of your enquiry here, along with our reply.
This helps us plan who picks up your enquiry and when.
We read every enquiry and reply within two working days.

What's included

Six areas of scope
01

A test dataset that reflects reality

Built from real inputs, including the difficult and ambiguous ones, labelled by people who know the subject and then frozen. A dataset chosen to make a system look good is worse than none at all, because it creates false confidence.

02

Scoring chosen before results are seen

Exact match, similarity of meaning, human rating or scoring by another model. Each is right somewhere and misleading elsewhere, and choosing after seeing the output is how evaluations turn into justifications.

03

Errors measured instead of asserted

For systems that answer from retrieved documents, the question is whether each answer is supported by the material it retrieved. That can be measured, and it should be reported as a number instead of as reassurance in a proposal.

04

Repeat testing as the system changes

Models are updated, prompts are edited and retrieval is tuned. Without a test you can run again, a decline in quality reaches you through your users.

05

Failure analysis

A close look at the cases the system gets wrong, grouped by type, which usually reveals a few specific problems to fix instead of a general need for a better model.

06

Human review where it matters

Clear rules for which outputs a person checks before they are used, based on the cost of a mistake instead of on convenience.

How the work runs

STEP 01

Scope the task

What the system must do, and what would count as failure.

STEP 02

Build the dataset

Real inputs, labelled by people who know the subject, with difficult cases included.

STEP 03

Agree the scoring

The method and the threshold, written down before any results exist.

STEP 04

Run, report and run again

Including after every later change to the system.

When to come to us

  1. 01You are being asked to approve an AI deployment on the strength of a demonstration.
  2. 02A system is already live and nobody can say whether it is getting better or worse.
  3. 03The output feeds a decision where being wrong has a real cost.
  4. 04Two providers are competing for the same work and you want them compared on the same test.

What we do not do

  • Evaluating on a dataset the builder assembled.
  • Reporting a single headline accuracy figure without showing how the difficult cases performed.
  • Publishing benchmark figures from someone else's system as though they were yours.

Common questions

Can we not just try it and see?
You can, and for low-risk internal tools that is often reasonable. Where the output feeds a decision that costs money to get wrong, trying it and seeing means learning the failure rate through the failures.
Who labels the test dataset?
People who know the subject, usually from your team instead of ours. Labelling is where subject knowledge enters the evaluation, and handing it to outsiders defeats the purpose.
How big does the dataset need to be?
Smaller than most people expect, as long as it is representative and includes the hard cases. A few hundred well-chosen examples are worth more than several thousand easy ones.
Can you evaluate a system we have already built?
Yes, and that is often when an evaluation is most useful. We test the system as it stands and report what we find, including the parts that work well.
How often should we repeat the evaluation?
After every meaningful change to the model, the prompts or the retrieval, and on a regular schedule even when nothing has changed, because the services behind these systems are updated frequently.
Enquire

Tell us what you need help with.

Share a few details below and we will reply within two working days. If it is not something we can help with, we will say so and, where we can, suggest someone who can.

Or send your enquiry here

It helps us understand your business before we reply.
We use it only to reply to this enquiry, usually on WhatsApp.
We send a copy of your enquiry here, along with our reply.
This helps us plan who picks up your enquiry and when.
We read every enquiry and reply within two working days.

Kaelo Global is a Dubai company licensed in Meydan Free Zone, with over 100 clients served so far.