
Where AI earns its place in an operating business, and where it does not. Part of our management consultancy.
What this activity does
Most AI work dies between the demonstration and the Monday morning it is supposed to run. We advise operators on the unglamorous half, meaning the evaluation, the test dataset, the step where a person checks the uncertain cases and the real cost at production volume. That half decides whether anyone is still using the system in six months, and it is the part a vendor has least reason to discuss.
The question we are usually asked is whether a particular task should be automated at all. Often the answer is that a simpler change would do more, such as fixing a form, removing a duplicate step or writing down a rule that currently lives in one person's head. We would rather say that in the first week than build something impressive that quietly stops being used.
Where the answer is yes, the work is mostly unglamorous. Getting the right documents in front of the model, deciding what happens when it is unsure, measuring how often it is wrong and in which direction, and keeping the cost per document or per case where the business case assumed it would be.
We do not sell software, and we have no licence income to protect. If an off-the-shelf product fits your situation, we will tell you which one and help you test it properly before you commit.
Tested on our own work first
Misclassified cloth is expensive twice, in duty paid and in time lost at the border. We built the classifier for our own import desk and ran it alongside the human process instead of in place of it, and we kept it there until its disagreements with the broker were the interesting ones. We only offer this kind of work to clients because it survived our own volume first.
How an engagement works
We establish the question the work has to answer, and write it down with the measure of success, before anything is bought or built.
The workflow and its measurement go live together, running alongside the existing process until the results are good enough to rely on.
The evaluation, the documentation and the routine for handling exceptions stay with your team, so the system can be maintained without us.
Where it is used
Common questions
Almost always the second. What decides whether an AI feature survives contact with an operating team is retrieval, evaluation and the handling of the cases the model gets wrong, instead of training a model from scratch. That layer is what we build.
Through an evaluation built from real examples your team has already judged. It runs before the system is used and after every change, so a decline in quality is caught by the test instead of by a customer.
It is planned for in advance. Every deployment has a defined route for uncertain cases, usually to a person, and every one is logged. A workflow with no answer for the wrong case is a workflow that gets switched off in the second month.
Yes, although we use things ourselves first. Anything we offer a client has run against our own freight, textile or online retail volume, which is why our recommendations come with an opinion about what will not work.
The cost of running it is part of the design instead of a surprise afterwards. The choice of model, caching and routing are all set against a stated cost for each transaction, because a task that is cheaper done by hand should be done by hand.
We agree at the start where your data is processed and stored, what a provider may do with it, and which material must never leave your own systems. Those answers are written into the engagement before any pilot begins.
Usually a few weeks. Most of that time goes into gathering real examples and agreeing how success will be measured, which is the part that makes everything afterwards quicker.
Begin
If AI is not the right tool for it, we will tell you.