Compatibility check
Can I run Llama 3.1 8B Instruct on Apple M1 (8-core GPU, 8GB unified)?
Probably partially, with CPU offload.best quant:
This one is close. The answer rests on a file size we calculated from the parameter count rather than measured. A 10% error in that size changes the verdict, so treat it as a maybe and test before you buy.
5 GB of weights, plus 1.5 GB for the software that runs it and the smallest conversation it can hold, comes to 6.5 GB against the 6 GB this 8 GB device leaves free.
Decode speed
—
Usable context
—
Memory
8 GB
Spills to system RAMest
Spills to system RAM
Too large
Plan B — rent it
Llama 3.1 8B Instruct is hosted by 6 providers from $0.080 per 1M output tokens (Groq).Compare providers →