AI Solutions · Service
What AI actually costs to run
Cost per token at production volume, and the exception queue nobody prices.
At a glance
What it is
“The model cost is usually the smallest line in the bill.”
A principal, Kaelo AI Solutions
What’s included
Four areas of scopeCost per unit of work, not per token
Tokens are the billing unit; the meaningful figure is cost per document, per query or per resolved case — including the retries and the failures that consumed tokens without producing an answer.
Context length as the hidden multiplier
Sending a large context on every call is the most common reason a business case is wrong by an order of magnitude. Retrieval quality is a cost decision as much as an accuracy one.
The exception queue, priced
Whatever proportion the system cannot handle confidently still needs a person. That is a real, ongoing cost and it belongs in the case from the start.
Build against buy, argued honestly
Sometimes an off-the-shelf product is cheaper than the engineering time to maintain a bespoke one. We have no licence revenue riding on the answer.
How the work runs
Volume model
Real expected throughput, including peaks.
Unit cost build
Per unit of work, with retries and context accounted.
Exception rate estimate
From the evaluation, not from optimism.
Total cost of ownership
Including engineering maintenance and model change risk.
When to come to us
- 01 You are approving an AI business case built on a pilot's usage pattern.
- 02 A live system's costs have grown faster than its usage.
- 03 You are choosing between building and buying.
What we do not do
- Publishing model pricing as though it were stable — it changes, and a page quoting it ages badly.
- Estimating exception rates before an evaluation exists.
Common questions