Models /GLM 4.5 /Can I run it?
Compatibility check

Can I run GLM 4.5 on B200 (SXM 192GB)?

Partially — with CPU offload.best quant:
Decode speed
—
Usable context
—
Memory
192 GB
Spills to system RAM
Spills to system RAMest
Too large
What is quantisation? →
Plan B — rent it

GLM 4.5 is hosted by 3 providers from $2.20 per 1M output tokens (Novita AI).Compare providers →

GLM 4.5
358B / 32B · full specs & benchmarks →
B200 (SXM 192GB)
192 GB · everything it runs →