Models /Qwen3 VL 235B A22B Thinking /Can I run it?
Compatibility check

Can I run Qwen3 VL 235B A22B Thinking on Apple M2 Ultra (76-core GPU)?

Probably partially, with CPU offload.best quant:

This one is close. The answer rests on a file size we calculated from the parameter count rather than measured. A 10% error in that size changes the verdict, so treat it as a maybe and test before you buy.

148.6 GB of weights, plus 6 GB for the software that runs it and the smallest conversation it can hold, comes to 154.6 GB against the 144 GB this 192 GB device leaves free.

Decode speed
—
Usable context
—
Memory
192 GB
Spills to system RAMest
Spills to system RAM
Too large
What is quantisation? →
Plan B — rent it

Qwen3 VL 235B A22B Thinking is hosted by 3 providers from $4.00 per 1M output tokens (Alibaba Cloud).Compare providers →

Qwen3 VL 235B A22B Thinking
236B / 22B · full specs & benchmarks →
Apple M2 Ultra (76-core GPU)
192 GB · everything it runs →