Models /Qwen2.5 VL 72B Instruct /Can I run it?
Compatibility check

Can I run Qwen2.5 VL 72B Instruct on Apple M4 Max (32-core GPU)?

Probably partially, with CPU offload.best quant:

This one is close. The answer rests on a file size we calculated from the parameter count rather than measured. A 10% error in that size changes the verdict, so treat it as a maybe and test before you buy.

46.3 GB of weights, plus 3.3 GB for the software that runs it and the smallest conversation it can hold, comes to 49.6 GB against the 48 GB this 64 GB device leaves free.

Decode speed
—
Usable context
—
Memory
64 GB
Spills to system RAMest
Spills to system RAM
Too large
What is quantisation? →
Plan B — rent it

Qwen2.5 VL 72B Instruct is hosted by 3 providers from $1.00 per 1M output tokens (OpenRouter).Compare providers →

Qwen2.5 VL 72B Instruct
73.4B · full specs & benchmarks →
Apple M4 Max (32-core GPU)
64 GB · everything it runs →