Compatibility check
Can I run Qwen3 Next 80B A3B Thinking on Apple M5 Pro (20-core GPU)?
Partially — with CPU offload.best quant:
51.3 GB of weights, plus 2.9 GB for the software that runs it and the smallest conversation it can hold, comes to 54.2 GB against the 48 GB this 64 GB device leaves free.
Decode speed
—
Usable context
—
Memory
64 GB
Spills to system RAM
Spills to system RAM
Too large
Plan B — rent it
Qwen3 Next 80B A3B Thinking is hosted by 4 providers from $1.20 per 1M output tokens (Google Vertex AI).Compare providers →