Running a model on your own machine sounds like the enthusiast's option, and for a good number of people that is exactly what it is. There are four situations where it is simply the right answer, and one where the honest advice is to keep renting.
The first is privacy you can verify. A downloaded model runs with the network cable out, figuratively or literally, and prompts that never leave your machine cannot be logged, trained on or subpoenaed, as a physical fact rather than a policy. If your work touches client data, patient notes or unreleased code, this is the strongest position available, and no provider's retention wording matches it. Everything on our provider pages about who can see what is an attempt to get close to this, from the other direction.
The second is cost, and since that is the claim a sceptic arrives to check, here is the sum instead of the slogan. Take the price of the card. Divide it by what you spent on models last month. That is your payback in months, and it is the only version of this argument worth repeating. Our card pages carry each card's list price, which is one half of it. Your own bill is the other half, and nobody else can supply that number for you.
Run it and the answer splits cleanly in two. If you pay for one flat-rate subscription, you are not spending enough for the sum to land inside a year, so the card has to earn its place on privacy, or offline, or the tinkering below, and in practice it usually can. If you carry a working bill instead, a small team, an agent that reruns overnight, a summarising job that never sleeps, then the division starts coming back in months instead of years, and the question worth your time becomes whether a model that fits the card does the work. Put your own two numbers in before you believe either answer. The two picks below are where to test the second half of it, and the second one's page shows you what the hosts charge to serve the very same model, so half the comparison is already done — while the fit check says which of them the machine on your desk already holds.
One hole in our half of the sum, since it changes the answer: we hold each card's list price and not its street or used price. The cards people actually buy for this are last-generation 24 GB cards from the second-hand market, and the used price is the number that decides it. You will have to bring that one yourself.
The third reason is working offline. On a plane, on a site with no signal, in a country where a provider blocks you by region, a local model does not notice.
The fourth is the freedom to take things apart. Fine-tuning on your own data, wiring a model into home automation, running a transcriber across years of voice notes: things that are awkward or expensive through somebody else's API become weekend projects when the model sits on your disk and answers to you.
Against all four stands capability. The strongest models are not downloadable, and a laptop-sized model is not one of them, so if the job needs the best answer available then rented frontier models win and it is not close. Plenty of people land on both, keeping a local model for the private and the routine and reaching for a rented one when the problem is hard. That is a sensible place to end up, and it costs less than either purist position.
Where next: Check what your own machine runs · Card prices and bandwidth · Who can see your prompts · What is quantisation