There are three ways to buy someone else's compute, and most people arrive having met two of them. You can pay by the token at a hosted endpoint, you can rent a whole graphics card by the hour, or you can take the rung in between, where a card wakes up when a request arrives and goes away again afterwards.
The order to consider them in is short enough to carry around. If the model you want is already served by the token, take that price and stop reading, because that fight is won. If you have your own model and use it in bursts, price the serverless rung. If you have your own model and steady heavy traffic, or a privacy requirement you can state in a sentence, rent the card and hire the patience to run it.
Paying by the token wins as often as it does because of arithmetic that happens on the provider's side. They pack thousands of customers onto each card, so the card is never idle and the cost of the idle time is nobody's. You get the finished price and never think about hardware again. The two picks below are the shape of it: one model that no household could hold and one that comfortably fits a single 24 GB card, both of them sitting on a shelf of competing hosts with the prices live on their pages.
Rent a card by the hour and that arithmetic becomes yours. The meter runs while the card works and it runs while the card waits, and inference waits most of the time, so a card you keep at a fifth utilisation is charging you five times its sticker rate per token. The software underneath is yours too: drivers, a serving framework, model updates, and whatever wakes up when it falls over at three in the morning. People price the card and forget they have also hired a job.
What the hourly card buys that nothing else does is control. A model you fine-tuned that nobody else serves. A retention story you can say in one sentence, single tenant, your logs, nobody else's policy. Latency you tune yourself. When one of those is a real requirement, the cost and the work are simply what it costs, and the thing to check first is that the requirement is real.
The middle rung is where this guide has to be honest with you. A serverless GPU gives you a card that starts when a request arrives, bills by the second, and disappears again, which for a custom model used in bursts is the obvious shape. We hold nothing on it. We track no platform on this rung, so we have no price to show you, no cold-start figure we measured, and no host we can point at and stand behind. That is a hole in our data rather than a small market, and it is ours to fill. Everything either side of it is here: per-token offers with daily prices on every model page, and the cards themselves, with their memory, bandwidth and list price, in the hardware index.
So price it yourself, and measure the part everyone guesses at. The cold start is your model file being read off storage and into the card's memory, which means it scales with the size of the file and with nothing else interesting. Each model page carries that size at the common quantisation, so a model needing 17 GB will always wake more slowly than one needing 5, on any platform, and you can rank candidates before you sign up for anything. Then take one afternoon: deploy the model, send a request, wait past the platform's idle timeout, send the same request again, and time both. The second number is the one your users will feel, and an afternoon of your own measurements beats any range somebody else quotes you.
Where next: Hosted models and prices · The hardware index · Who can see your prompts