Concept

What is meant by "providers"

Companies that run models on their hardware and sell access by the token.

What is meant by "providers"

A provider runs models on its own hardware and sells you access, usually priced per million tokens. The confusion worth clearing up first: the company that made a model and the companies that serve it are often different, and for downloadable models there is real competition — the same Llama or Qwen weights might be served by ten providers at ten different prices.

Three kinds of company get called a provider, and they behave differently:

The labs themselves. OpenAI, Anthropic and Google sell access to their own models. For models that cannot be downloaded, the lab is usually the only door, and the lab's own price is the list price.

Independent hosts. Companies like DeepInfra, Together and Fireworks take downloadable models and serve them, competing on price and speed. This is where the same model gets ten prices — and where checking the provider table on any model page can cut your bill several-fold for identical output.

Routers. OpenRouter and similar sit in front of many hosts behind one API key, choosing where each request lands. Convenient, one bill, one integration — at the cost of one more party handling your traffic.

Two things differ between providers besides price, and both are on every provider page here. Speed — the same model genuinely runs at different speeds on different hosts, and we show measured throughput where we have it. And what happens to your prompts: whether a provider trains on them, logs them, or offers zero data retention. Where we do not know, the page says unknown — an absent answer is not a "no".

Where next: Providers, priced and compared · What is inference

Terms this page anchors

providerendpointzero data retention

Updated 2 Aug 2026