How to · Building with models

Estimating what an AI feature will cost before you build it

Work one realistic call all the way through and everything after it is multiplication. What catches people out is that most of a feature's bill is text nobody typed: instructions, retrieved context, conversation history and tool definitions, re-sent in full on every single call.

Updated 24 Sept 2026

Somebody always asks what the feature will cost before anyone has written it, and the honest answer is available the same afternoon. The bill is arithmetic: tokens in times the input rate, plus tokens out times the output rate, times the number of calls. Estimates rarely come apart at the multiplication. They come apart in three places where people count the wrong thing.

Take a support-triage feature and count it properly. The system prompt carries the instructions, six category definitions and four worked examples, and runs to around 1,200 tokens. Two similar past tickets are retrieved and pasted in for context, another 600. The ticket itself, a longish email, is 300. So the model reads about 2,100 tokens to produce a category and a one-line reason, which comes back at roughly 40. At 2,000 tickets a day that is 4.2 million tokens in and 80,000 out, and across a month, 126 million in against 2.4 million out.

Sit with that ratio for a second, because it is the whole first lesson. Output is under two per cent of the traffic, and output is the part priced several times higher. Anyone who estimated this feature from the ticket alone, the only text a human actually wrote, undercounted the input sevenfold. Measure a realistic call end to end.

The second place estimates break is the call count. A triage step that looks something up, reads what came back and tries again is one feature making five calls, and each of those calls re-reads everything before it, so the five-call version costs more than five times the one-call version. Count the calls along your worst realistic path.

The third is pricing a model you will not use. The spread between models of similar ability is large and it moves weekly, which is what the cheapest price anyone charges for each level of ability, tracked daily exists to follow, and the same open model is served at quite different rates depending on the host. Price your actual candidate on its actual host, and while you are there note the levers that host offers: caching for the repeated opening of your prompt, a batch endpoint for work nobody is waiting on, and a spend cap per key as the backstop.

The last multiplication is the one this page deliberately leaves to you, because the two numbers it needs move week by week and a figure typed here would be wrong before you read it. They sit on the cards below, live: a rate for every million tokens read, a rate for every million written. Take your 126 million to the first and your 2.4 million to the second, add them together, then double the total for everything you have not thought of yet.

Then there is the structural choice, which changes the slope of the bill. For work that is narrow and repetitive, sorting and extracting and tagging, a small model usually clears the bar, and the gap against a frontier model is a different order of bill. Price both. The comparison takes twenty minutes and it is the only one on this page that can change the answer by a factor rather than a percentage.

Where next: What is a token · Compare model prices · Using a small model as a classifier

What to run it with

Questions people actually ask

Input or output — which one should I be watching?+

Input, in most production features, even though output is priced several times higher per token. A feature that reads a long prompt and writes a short answer spends almost all of its tokens reading. The exceptions are real and easy to spot: anything that drafts, summarises at length or writes code produces enough output to tip the balance back.

Does prompt caching actually make a difference?+

Where a host offers it and your prompt has a large fixed opening, yes, and that is exactly the shape most features have. The repeated part is charged at a reduced rate on later calls. It varies by host and by model, so confirm it on your own candidate before you count on it, and keep the uncached figure as your planning number.

What if nobody can tell me the volume yet?+

Price a thousand calls and hand that number over. "Roughly this much per thousand triages" is a figure a product owner can reason about and a per-month guess is not, and it survives the volume being revised three times before launch.

Should I just take the cheapest host?+

Take the one you would be content to explain later. The same open model is served at very different rates, and the difference sometimes reflects a quieter machine or a shorter served context. Check who the host is and what they do with your prompts before a feature depends on them.

How far out is a careful estimate, usually?+

Far enough that the doubling at the end of the method earns its place. Retries, a longer system prompt than the one you drafted, a second call nobody counted, and traffic that arrives in bursts all push the same direction. An estimate that comes in under the real bill is the normal failure, which is why the method ends with a margin.

Twenty realistic calls, the median counts in and out, times your daily volume, times the two live rates, doubled. It is an afternoon's work, and it is the difference between a feature that ships and one quietly switched off in month three.