About LLMap

Why you're here

You've been using ChatGPT or Claude the way most people do now: a daily thought partner, a draft checker, a way to get unstuck. Maybe you've gone further and pointed Claude Code at an API, wired up a couple of agents, let them run overnight, and discovered how fast tokens add up when a model works while you sleep. It works, until something shifts. Maybe the bill spiked. Maybe pasting your company's code into someone else's cloud started to feel wrong. Maybe you read that someone is running a capable coding model on their own laptop, paying almost nothing, and thought: wait, can I do that?

So you looked around. You opened Hugging Face, which is wonderful if you read equations for breakfast and a wall of numbers if you don't. You found a benchmark site that says Model A scored 78.2 and Model B scored 79.1, both at something called q4_K_M, and learned nothing except that you're not the audience for that page. You tried Reddit, where three people gave three opposite answers with total confidence. Depending on your machine and what "fit" means, all three might be right.

The tools are out there. The judgment isn't.

That's the gap this exists to fill. We're not trying to out-catalogue Hugging Face, out-chart the analysts, or become another forum full of opinions. We're a translation layer for people like you: smart enough to learn this stuff, busy enough to want it explained without condescension and without the hype.

From "what can I run?" to "it's running"

Tell us what you have: a MacBook with 16GB of RAM, a gaming laptop, a monthly budget, or just curiosity. We'll show you which models actually fit and explain the trade-offs in plain language, including what a smaller, compressed version of a model gives up, what it still does well, and whether that matters for the chatting, coding or agent-running you have in mind. Want to stay local? We'll point you to the simplest runner and the exact settings. Want to try hosted? We'll show you the cheapest credible provider and what "cheaper" really means once you count tokens. Not sure of the difference yet? We'll start there.

Privacy is part of the answer

For a lot of people, privacy is the real reason to look local in the first place: a model running on your own machine never sends your code, your documents, or your questions anywhere. Hosted inference can still be the right call, but it deserves open eyes. Some providers train on your prompts. Retention policies range from zero to indefinite. And a US-based provider can be legally compelled to hand over customer data, wherever the server happens to sit. We record those policies where a provider publishes them: who trains on what, what gets logged, and where it runs. We do not have a complete answer for every provider yet, and where we don't, the page says so rather than leaving a gap you might read as safe. That way, "cheapest" never quietly means "least private".

How we keep it honest

Every recommendation is anchored to live data: current prices, real hardware fits, providers that are still around. The editorial take tells you what the numbers mean; the numbers keep the editorial take honest. When we're not sure, we say so. When our pick changes, we say why.

Sizes and fit verdicts say whether they are measured or estimated, because a figure we worked out should not wear the authority of one we checked. Where we don't hold something yet, the page says that plainly instead of leaving a blank you might read as a zero. A missing privacy answer is the one that matters most, so it never renders as reassuring.

How to use this site

Three ways in. Any of them gets you to the same place: a model you can actually run, and a sense of what to expect from it.

Stuck on a word? The guides explain the concepts one at a time, in the order they actually come up. And you can ask a question on the home page in plain language, then see the data behind the answer.

The whole trip, from your machine to a model to actually running it and knowing what to expect, takes one session. You don't need to become a machine-learning engineer. You need curiosity, a machine or a credit card, and about fifteen minutes.

That's why you're here.