Deciding · Running cheaper

Leaving a flat-rate chat subscription for a metered API

A monthly chat subscription is a good deal for heavy use of one lab's models. Paying per token is usually gentler than it sounds, and the worry has a cure that takes five minutes: count a heavy month properly, then set a spend cap at whatever you were already handing over.

Updated 22 Sept 2026

The monthly subscription is how most people meet this, and for heavy use of one lab's models it is a good deal. Paying per token instead feels like the risky choice, because an uncapped bill is a genuinely unpleasant idea to sit with. What the key buys in return is choice: any model on the market, the same key working in your editor and your tools as well as a chat window, and a quiet month that costs you nothing at all.

The reason the fear survives contact with the prices is that the unit is unfamiliar. Rates are quoted per million tokens, and a million tokens is a great deal of text. A paragraph runs to about a hundred. So the useful thing to do once, carefully, is count a heavy month.

Say a reply comes back at a few paragraphs, call it 400 tokens, and your question is 60. A twenty-turn conversation therefore has around 8,000 tokens written back to you. The reading is the larger and much less obvious half, because a chat app re-sends the whole conversation on every turn: by the twentieth question the model is reading the previous nineteen exchanges again, and the conversation totals somewhere close to 90,000 tokens read. Three conversations like that a day, five days a week, four weeks, and a heavy month lands at roughly five million tokens read and half a million written.

Both rates sit on the cards below, live, quoted per million. Five of the first plus half of the second is your month. Put that beside what the subscription takes from you each month and the comparison is one you have made yourself, which is the only version of it worth trusting.

Cap it in either case. Every serious provider lets you set a monthly spend limit, and a router gives you one limit across every model behind it. Set the cap to what the subscription used to cost and the worst case is that you have paid what you were already paying.

The honest ledger of what you give up is mostly about the app rather than the model. A subscription bundles memory of past chats, file handling, voice, image generation and image understanding, and a mobile experience somebody has polished for years. Going metered means choosing a chat front-end that takes an API key, and the good ones are genuinely good, but it is a setup step with an afternoon in it. Some people land on both, keeping a cheap subscription for the comforts and using the key for everything else, and that is a real answer.

The switch itself is three moves. Make a key at a provider or a router, set the spend cap, then point a chat app or your coding tools at it. Most of the afternoon goes on choosing your first model, which is what the verdicts on each model page are there to shorten.

Where next: What is a token · What is a model router · Compare model prices

What to run it with

Questions people actually ask

What does a genuinely heavy month look like?+

Three long conversations a day, five days a week, four weeks, is about five million tokens read and half a million written. That is a person using this as a working tool, and it is the figure to price against. If your use is lighter than that, and most people's is, the arithmetic only moves in your favour.

Why is the reading figure so much larger than the writing one?+

Because a chat app re-sends the whole conversation on every turn. The model has no memory between messages, so the transcript goes back with each new question, which means a long conversation is read many times over and written only once. It is the single thing people leave out when they guess at their own usage.

What actually happens when I hit the cap?+

Calls stop being served until you raise it or the month turns over. That is the whole mechanism, and it is why the cap is the thing to set first: it converts an open-ended worry into a known maximum. A router puts one cap across every model behind it, which is easier to hold in your head than four.

Do I lose my chat history?+

The history lives with whichever front-end you use, so it moves when you move. Exports are worth taking before you cancel anything. What genuinely does not travel is a subscription's memory of past chats: that is a feature of the app you are leaving, and a new front-end starts you at nothing.

Is one key enough?+

One is enough to start, and a router is the version of that answer which does not box you in, because it gives you many models and one bill behind a single key. Plenty of people run two keys instead, one at a lab and one at a host serving open models, and split the work by what each is good at.

Set the cap to what the subscription used to cost, and the worst case is that you have paid what you were already paying.