Deciding · Choosing hardware

Renting a GPU in the cloud: three ways to buy someone else's compute

Almost everybody who asks about renting a GPU turns out to want something cheaper and duller than a GPU. The expensive answer is famous, the right answer usually is not, and the difference between them is a bill that runs all night for work that took four minutes.

Updated 24 Sept 2026

There are three ways to buy someone else's compute, and most people arrive having met two of them. You can pay by the token at a hosted endpoint, you can rent a whole graphics card by the hour, or you can take the rung in between, where a card wakes up when a request arrives and goes away again afterwards.

The order to consider them in is short enough to carry around. If the model you want is already served by the token, take that price and stop reading, because that fight is won. If you have your own model and use it in bursts, price the serverless rung. If you have your own model and steady heavy traffic, or a privacy requirement you can state in a sentence, rent the card and hire the patience to run it.

Paying by the token wins as often as it does because of arithmetic that happens on the provider's side. They pack thousands of customers onto each card, so the card is never idle and the cost of the idle time is nobody's. You get the finished price and never think about hardware again. The two picks below are the shape of it: one model that no household could hold and one that comfortably fits a single 24 GB card, both of them sitting on a shelf of competing hosts with the prices live on their pages.

Rent a card by the hour and that arithmetic becomes yours. The meter runs while the card works and it runs while the card waits, and inference waits most of the time, so a card you keep at a fifth utilisation is charging you five times its sticker rate per token. The software underneath is yours too: drivers, a serving framework, model updates, and whatever wakes up when it falls over at three in the morning. People price the card and forget they have also hired a job.

What the hourly card buys that nothing else does is control. A model you fine-tuned that nobody else serves. A retention story you can say in one sentence, single tenant, your logs, nobody else's policy. Latency you tune yourself. When one of those is a real requirement, the cost and the work are simply what it costs, and the thing to check first is that the requirement is real.

The middle rung is where this guide has to be honest with you. A serverless GPU gives you a card that starts when a request arrives, bills by the second, and disappears again, which for a custom model used in bursts is the obvious shape. We hold nothing on it. We track no platform on this rung, so we have no price to show you, no cold-start figure we measured, and no host we can point at and stand behind. That is a hole in our data rather than a small market, and it is ours to fill. Everything either side of it is here: per-token offers with daily prices on every model page, and the cards themselves, with their memory, bandwidth and list price, in the hardware index.

So price it yourself, and measure the part everyone guesses at. The cold start is your model file being read off storage and into the card's memory, which means it scales with the size of the file and with nothing else interesting. Each model page carries that size at the common quantisation, so a model needing 17 GB will always wake more slowly than one needing 5, on any platform, and you can rank candidates before you sign up for anything. Then take one afternoon: deploy the model, send a request, wait past the platform's idle timeout, send the same request again, and time both. The second number is the one your users will feel, and an afternoon of your own measurements beats any range somebody else quotes you.

Where next: Hosted models and prices · The hardware index · Who can see your prompts

What to run it with

Questions people actually ask

Is the hourly card ever the cheap option?+

When you can keep it genuinely busy. The bill runs whether the card is working or idle, and an inference workload is idle most of the time, so the effective price per token climbs by however much of the hour you wasted. A card at a fifth utilisation is five times its sticker rate.

What does serverless GPU cost, and how slow is the first request?+

We do not hold either figure. We track no serverless platform at all, so we have no price, no cold-start measurement and no host we can name, and we would rather say that than hand you a number that came from nowhere. The last section of this guide is how to measure the cold start on your own model in an afternoon.

Do I need a GPU at all to serve a model to a few people?+

Often not. A desk machine with enough memory serves a household or a small team perfectly well, and the fit checker will tell you which models it holds. Renting starts to make sense when the traffic is other people's and it has to be there at three in the morning.

Two of these three rungs are in the archive with live prices against them. The middle one you will have to price yourself, and the last section tells you how to measure the part everyone guesses at.