Models / Best of /local-32gb
Updated October 2026 · newest first, from live data

Best models for 24–32 GB VRAM

Open-weight models needing roughly 10–28 GB at Q4, newest first. Check the size column against your memory: a 24 GB card (3090/4090) covers the lower rows; a 32 GB machine covers the range. Newest first, not a ranking by quality. 3 of the 10 below carry a rating from a board we track, and the highest placed is Qwen3.8 27B (3/5 everyday (Arena Text (overall), 63 of 168)). For the other 7 we hold no score at all, which is not the same as a low one. Every row links to full specs, hardware verdicts and provider pricing.

What is quantisation? →
1Ternary Bonsai 2 27Bprism-mlno rating from any board we trackest27B$0.075 / $0.50per 1M · via OpenRouter2Nex-N2.5-MiniNex AGIno rating from any board we trackest35.1B$0.025 / $0.10per 1M · via OpenRouter3Muse Glimmer 30BMeta3/5 everyday (Arena Text (overall), 75 of 168)est29.8B$0.30 / $1.10per 1M · via Phala4Qwen3.8 27BQwen3/5 everyday (Arena Text (overall), 63 of 168)est27.8B$0.15 / $1.88per 1M · via Phala5Nemotron 3.5 LightningNVIDIA2/5 everyday (Arena Text (overall), 125 of 168)est31.6B$0.059 / $0.17per 1M · via Io Net6Laguna XS 2.1poolsideno rating from any board we trackmeasured33.4B$0.060 / $0.12per 1M · via OpenRouter7Nex-N2-MiniNex AGIno rating from any board we trackest35.1B$0.025 / $0.10per 1M · via Nex AGI8Hy-MT2-30B-A3BTencentno rating from any board we trackest30.1B / 3B$0.074 / $0.29per 1M · via OpenRouter9Qwen3.6 27BQwenno rating from any board we trackest27.8B$0.32 / $2.70per 1M · via Phala10Qwen3.6 35B A3BQwenno rating from any board we trackest36B / 3B$0.15 / $1.00per 1M · via OpenRouter
More listsBest models for codingBest models for writingCheapest hosted modelsBest models that fit in 12 GB VRAMLongest context windowsBest open-weight modelsFastest hosted models