Best local AI for Apple M3 Max · 36GB

The best local AI on a Apple M3 Max · 36GB right now is Qwen3 (30B-A3B (MoE) @ Q4_K_M, ~65–110 tok/s). Full ranking below — capability × real speed on this hardware, never sponsored.

#modelbest fitverdictspeed
1 Qwen3 30B-A3B (MoE) Q4_K_M ✅ Runs (tight) 65–110 tok/s
2 DeepSeek-R1 (distill) 14B Q5_K_M ✅ Runs comfortably 16–25 tok/s
3 Qwen2.5 14B Q5_K_M ✅ Runs comfortably 16–25 tok/s
4 Gemma 4 12B Q5_K_M ✅ Runs comfortably 16–25 tok/s
5 Gemma 3 12B Q5_K_M ✅ Runs comfortably 16–25 tok/s
6 Mistral Small 3 24B Q4_K_M ✅ Runs comfortably 12–20 tok/s
7 Llama 3.1 8B Q8_0 ✅ Runs comfortably 20–35 tok/s
8 Phi-4 14B Q5_K_M ✅ Runs comfortably 15–25 tok/s
9 Gemma 2 27B Q3_K_M ✅ Runs comfortably 12–19 tok/s
10 Mistral Nemo 12B Q6_K ✅ Runs comfortably 17–30 tok/s
11 Mistral 7B 7B Q8_0 ✅ Runs comfortably 20–35 tok/s
12 Llama 3.2 3B Q8_0 ✅ Runs comfortably 45–70 tok/s
$npx runlocal-sh install qwen3

Rankings use the same GQA-aware fit + bandwidth-calibrated speed math as the CLI — no sponsorships, no affiliate re-sorting, ever.

Different machine? Apple M1 · 8GB · Apple M1 Pro · 16GB · Apple M2 · 16GB · Apple M3 Pro · 18GB · Apple M2 Pro · 32GB · Apple M4 Pro · 48GB

Or check a specific model on your exact rig: open the checker →