Troubleshooting · Getting started

Why your local model is slow, and what to change

Below about ten tokens a second, waiting for an answer feels like waiting; above fifty it feels instant. If yours crawls, the cause is usually a model that very nearly fits and is quietly spilling into ordinary memory, where it runs at a fraction of the speed with nothing on screen to say so.

Updated 23 Sept 2026

Your model runs and the words crawl out. Before buying anything, it is worth knowing that speed here obeys one law: producing a single token means streaming essentially the whole model through the processor, so your ceiling is memory bandwidth divided by model size. A machine with twice the bandwidth is roughly twice as fast on the same model, and a model half the size is roughly twice as fast on the same machine. Every lever below moves one of those two numbers.

You can put a figure on your own ceiling in about a minute. The hardware index lists the memory bandwidth of each machine, and every model page lists the size of the file at each compression level. Bandwidth divided by file size is an upper bound you will not beat, and real output lands some way under it, but it tells you quickly whether you are chasing a small improvement or the wrong machine entirely.

The usual culprit is more specific than that, and it is worth ruling out first because it is invisible. When a model very nearly fits in graphics memory, the tools quietly leave the remainder in ordinary memory and carry on. Speed collapses, often to a tenth, and everything still works: no error, no warning, just a model that has become unusable while looking perfectly healthy. The answer is to make it fit. A smaller compression of the same model that lives entirely in graphics memory beats a larger one that spills, by far more than the difference in quality between them costs you. Check your machine against the model before you conclude anything about either.

Context is the lever people forget. The working memory for a conversation grows turn by turn, and past a point it eats the headroom the model needed and slows every step after. A chat that has got slower the longer it ran is doing exactly what it appears to be doing, so start a fresh one, or trim what the tool re-sends each turn.

Stepping down a model size is the lever people resist, and it is the one that most often works. Quality per parameter has improved quickly enough that this year's mid-sized models do work last year's large ones did, at several times the speed, and the ratings on each model page will tell you what a step down actually costs you before you assume you need the big one. One family bends the law outright: a mixture-of-experts design stores a great many parameters and puts only a few billion of them to work on each token, so it answers at closer to a small model's pace while still asking for a large model's memory. Where speed is your constraint and memory is not, that trade is a good one.

If the model fits, the context is short, the compression is sensible and it is still too slow, then the bandwidth law is what you are hearing. Bandwidth is the specification to shop on for this work, and the hardware index leads with the figure that spec sheets bury, per platform and per machine.

Where next: Check your machine · VRAM and what fits · What compression does to a model

What to run it with

Questions people actually ask

How do I tell whether my model is spilling?+

Most tools report how many layers went to the graphics card and how many stayed in ordinary memory, usually in the log line nobody reads on startup. Watch your graphics memory as the model loads: if it fills and the model keeps loading, the remainder has gone somewhere slower. The fit checker answers the same question before you download anything.

What counts as a reasonable tokens-per-second figure?+

Ten is about the point where it stops feeling like waiting, and fifty or so feels instantaneous, which is roughly the speed a person reads. Between those two it comes down to what you are doing: a long piece of writing is fine at fifteen, and a coding assistant that has to keep up with you is not.

Would more system memory fix it?+

It buys you the ability to run a model that spills, at the speed of a model that spills. System memory is the slow place the overflow goes, so adding more lets a larger model load and does little for how fast it answers. Graphics memory, or a smaller model, is what moves the number.

Does a long context slow things down on its own?+

Yes, twice over. The working memory for a conversation grows with every turn, which eats the headroom the model needed to fit, and every new token has more to look back over. A conversation that started fast and has become sluggish is usually telling you it is time to begin a fresh one.

Is a mixture-of-experts model simply faster, then?+

Faster for its size, and only that. It still has to be held in memory whole, so it does nothing for whether a model fits your machine. What it changes is the speed once it does fit, by reading a fraction of itself for each token instead of all of it.

Fit first, context second, model size third, hardware last. A model that spills out of graphics memory runs at a fraction of its speed and never mentions it.