What hardware runs a local local chat workflow?

A general chat model for everyday local conversation. The smallest cataloged machine that runs Gemma 3 is a Apple M1 · 8GB. Ranked below, smallest memory first — each model at its best quant, one model loaded at a time.

machineGemma 3
Apple M1 · 8GB 1B @ Q8_0 · ~35–55 tok/s
Apple M2 · 8GB 1B @ Q8_0 · ~50–80 tok/s
Apple M3 · 8GB 1B @ Q8_0 · ~50–80 tok/s
AMD Radeon RX 7600 · 8GB VRAM 4B @ Q8_0 · ~20–35 tok/s
NVIDIA GeForce RTX 3070 · 8GB VRAM 4B @ Q8_0 · ~30–55 tok/s
NVIDIA GeForce RTX 4060 · 8GB VRAM 4B @ Q8_0 · ~20–35 tok/s
NVIDIA GeForce RTX 3080 · 10GB VRAM 4B @ Q8_0 · ~60–95 tok/s
AMD Radeon RX 6700 XT · 12GB VRAM 12B @ Q4_K_M · ~14–25 tok/s
AMD Radeon RX 7700 XT · 12GB VRAM 12B @ Q4_K_M · ~16–25 tok/s
NVIDIA GeForce RTX 3060 · 12GB VRAM 12B @ Q4_K_M · ~14–25 tok/s
NVIDIA GeForce RTX 4070 · 12GB VRAM 12B @ Q4_K_M · ~20–35 tok/s
Apple M1 · 16GB 4B @ Q8_0 · ~7–12 tok/s

Fits computed by the same engine as the CLI — reproduce this table with:

$npx runlocal-sh advise gemma3

Already have a machine? Check your exact rig →