Compatibility check
Can I run Qwen3 Next 80B A3B Thinking on L40S?
Partially — with CPU offload.best quant:
Decode speed
—
Usable context
—
Memory
48 GB
Spills to system RAM
Spills to system RAM
Too large
Plan B — rent it
Qwen3 Next 80B A3B Thinking is hosted by 4 providers from $1.20 per 1M output tokens (Google Vertex AI).Compare providers →