How to · Building with models

You have tried one lab's API — how to try the rest

You have a key from one lab and a suspicion there is more out there. For work that does not need a frontier model, the same request answered by a different company costs a fraction of what you are paying now, and moving your code across is usually an afternoon.

Updated 23 Sept 2026

The shape nearly everyone implements is OpenAI's chat completions endpoint: a base URL, a key, a list of messages, a model name. If your code speaks that, pointing it at another company is two lines of configuration and an account. If it speaks a lab's own SDK, budget an afternoon for a small translation, because the message and tool formats differ even where the idea behind them is identical.

What that buys you is worth seeing for yourself. Put a lab flagship next to an open-weights model of the same generation and the comparison does the division for you, at today's prices, from the cheapest host tracked for each. The gap is large enough to change what you are willing to build, which is the whole reason this page exists.

Three things genuinely differ once you start moving around, and price is only the first. The second, and the one that catches people out, is how much room a host actually leaves you for your text — its served context. A host may run a long-context model with a smaller window to save memory, so the same model name can mean a different amount of room depending on where you bought it. Speed is the third. Some hosts serve on hardware that streams several times faster than others, which changes what an interactive product feels like to use.

All of that sits in one table on each model page. Gemma 4 31B's host list is a good first look, because it is one downloadable model served by a long list of companies, and those rows are not copies of one another. Some hosts run the full-size file and some serve a squeezed copy of it, the room they leave you for your text differs from row to row, and the prices differ again on top. Read across a row before you read down the price column.

One caution as you spread out. More providers means more parties handling your prompts, and their policies vary far more than their APIs do. Where a host has said nothing about training on your data, the table says unknown, not the answer you would prefer, and a blank is never a promise. Read the training column before the price column for anything you would mind seeing again.

Where next: Providers, compared · Who can see your prompts

What to run it with

Questions people actually ask

Will my existing code actually work elsewhere?+

If it speaks the common request shape — OpenAI's chat completions, which nearly every host implements — then yes: a changed base URL and a new key. If it speaks a lab's own SDK, budget for a small translation, because message and tool formats differ even where the idea behind them is identical.

What genuinely differs between providers serving the same model?+

Price, the compression the host runs, how much room it leaves for your text, its measured speed, and what it says about training on your prompts. Every model page puts them in one table, host by host, and where a host has stated nothing the table says unknown and does not assume the answer you would like.

Is the cheapest host the right host?+

Often, and not always. The same weights can be served with a shorter context window, on slower hardware, or under a policy that reserves the right to train on what you send. Read those three columns before the price column, and keep the cheapest row for work where none of them matters to you.

One request shape, many shops. Trying the rest is a changed base URL and a new key, and the gap between a lab's own price and a host's is wider than most people guess.