The two tools worth your time are LM Studio and Ollama, and they differ in temperament
more than in what they can do. LM Studio looks like a chat app: you browse models inside
it, click download and start typing, and it shows you which ones fit before you commit to
the download. Ollama is something you type — ollama run and a model name — and its real
value arrives later, because it quietly serves an API on your own machine, so editors,
scripts and coding agents can talk to your model
instead of a hosted one. So: the app if you want to be chatting in fifteen minutes, the
command if you expect to wire something onto it later.
Model names carry a number with a B after it, and it is worth decoding once. 9B means nine billion parameters, the internal numbers the model learned while it was trained. More of them usually means a better answer and always means a bigger file, and that file has to fit in your memory for the model to answer at a sensible speed. So the useful question is never which model is best; it is which of the good ones your machine holds.
The fit check answers that, and it is the only place on this site you need for it. Pick your machine from the list, or paste what your system reports about memory, and every model comes back with a verdict — it runs, it spills over into slower memory, or it will not fit — with the memory figure and an estimated speed beside it. Run that before you download anything, because a model that nearly fits is the worst outcome on offer: everything keeps working, just miserably.
One rule of thumb while you learn the shape of it: nine billion parameters sit inside 16 GB, and every step up in parameters asks for a step up in memory. Treat it as rough, because two things bend it. Compression is the first, since most people run a shrunk copy of a model, which is what makes any of this fit at all. The second is the mixture designs, where only a slice of the model works on each word: those need less memory than their name suggests. The picks below are three rungs of that ladder with their real verdicts one click away, and read the speed estimate beside the verdict as well — on an older laptop a model that fits comfortably can still answer at a pace you notice.
Expect the first answer to feel slower than a hosted service. That is your hardware being honest with you, and the fit check tells you roughly how slow before you spend the download. What you get in exchange: nothing you type leaves the machine, ever, and nobody is counting.
Where next: Check what your machine runs · Point your coding tools at it