The best local AI on a Apple M2 Ultra · 192GB right now is DeepSeek-R1 (distill) (70B (Llama) @ Q2_K, ~9–14 tok/s). Full ranking below — capability × real speed on this hardware, never sponsored.
| # | model | best fit | verdict | speed |
|---|---|---|---|---|
| 1 | DeepSeek-R1 (distill) | 70B (Llama) Q2_K | ✅ Runs comfortably | 9–14 tok/s |
| 2 | Qwen3 | 30B-A3B (MoE) Q8_0 | ✅ Runs comfortably | 60–95 tok/s |
| 3 | Qwen2.5 | 72B Q3_K_M | ✅ Runs comfortably | 7–11 tok/s |
| 4 | Llama 3.1 | 70B Q3_K_M | ✅ Runs comfortably | 7–11 tok/s |
| 5 | Llama 3.3 | 70B Q3_K_M | ✅ Runs comfortably | 7–11 tok/s |
| 6 | Mistral Small 3 | 24B Q4_K_M | ✅ Runs comfortably | 16–25 tok/s |
| 7 | Mixtral 8x7B | 8x7B (MoE) Q8_0 | ✅ Runs comfortably | 17–30 tok/s |
| 8 | Gemma 2 | 27B Q3_K_M | ✅ Runs comfortably | 15–25 tok/s |
| 9 | Gemma 4 | 12B Q8_0 | ✅ Runs comfortably | 15–25 tok/s |
| 10 | Gemma 3 | 27B Q8_0 | ✅ Runs comfortably | 8–13 tok/s |
| 11 | Phi-4 | 14B Q5_K_M | ✅ Runs comfortably | 20–35 tok/s |
| 12 | Mistral 7B | 7B Q5_K_M | ✅ Runs comfortably | 40–70 tok/s |
Rankings use the same GQA-aware fit + bandwidth-calibrated speed math as the CLI — no sponsorships, no affiliate re-sorting, ever.
Different machine? Apple M1 · 8GB · Apple M1 Pro · 16GB · Apple M2 · 16GB · Apple M3 Pro · 18GB · Apple M2 Pro · 32GB · Apple M4 Pro · 48GB
Or check a specific model on your exact rig: open the checker →