New in the catalog

Newest additions first — each with the cheapest cataloged machine that runs it. Dates are when a model entered the runlocal catalog. Follow along: RSS feed · @runlocal_sh

2026-07-07

2026-06-28

2026-06-26

2026-06-24

Code Llama 7B · 13B · 34B
Meta's Llama-2-based code models. Still popular on Ollama for completion.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (Q5_K_M)
Codestral 22B
Mistral's 22B code model with strong fill-in-the-middle. 32K context.
runs on a NVIDIA GeForce RTX 5080 · 16GB VRAM (Q3_K_M)
Command R 35B
Cohere's RAG/tool-use specialist. No GQA, so KV cache is heavy at long context.
runs on a NVIDIA GeForce RTX 5090 · 32GB VRAM (Q3_K_M)
DeepSeek-R1 (distill) 1.5B · 7B · 8B (Llama) · 14B · 32B · 70B (Llama)
Open reasoning model. Distilled into Qwen/Llama backbones you can actually run locally.
runs on a Apple M1 · 8GB (Q8_0)
Gemma 2 2B · 9B · 27B
Google's efficient open models. The 9B is a standout for its size. 8K context.
runs on a Apple M1 · 8GB (Q4_K_M)
Gemma 3 1B · 4B · 12B · 27B
Google's latest, with vision and 128K context. Strong quality per parameter.
runs on a Apple M1 · 8GB (Q8_0)
Llama 3.1 8B · 70B
Meta's workhorse general model. Reliable all-rounder with 128K context.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (Q8_0)
Llama 3.2 1B · 3B
Tiny Llamas for edge + on-device. Fast, light, 128K context.
runs on a Apple M1 · 8GB (Q8_0)
Llama 3.2 Vision 11B
Multimodal Llama — image understanding plus chat. 11B fits a 16GB+ Mac.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (Q4_K_M)
Llama 3.3 70B
70B that rivals far larger models. The best open general model that fits a 64GB Mac.
runs on a NVIDIA GeForce RTX 5090 · 32GB VRAM (Q2_K)
LLaVA 7B · 13B
Popular open vision-language model for image Q&A and captioning.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (Q5_K_M)
Mistral 7B 7B
The classic fast 7B. Apache-licensed, runs anywhere.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (Q8_0)
Mistral Nemo 12B
12B with 128K context, built with NVIDIA. Great mid-size all-rounder.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (Q5_K_M)
Mistral Small 3 24B
24B that punches at 70B level with much faster generation. 32K context.
runs on a NVIDIA GeForce RTX 5080 · 16GB VRAM (Q3_K_M)
Mixtral 8x7B 8x7B (MoE)
Sparse MoE: 47B total but only ~13B active per token, so it decodes fast — if it fits in memory.
runs on a NVIDIA GeForce RTX 5090 · 32GB VRAM (Q4_K_M)
mxbai-embed-large 335M
Strong open embedding model for retrieval. Small and quick.
runs on a Apple M1 · 8GB (F16)
Nomic Embed Text 137M
Fast, high-quality text embeddings for RAG. Tiny — runs anywhere.
runs on a Apple M1 · 8GB (F16)
Phi-3.5-mini 3.8B
3.8B with 128K context. No GQA, so KV grows fast at long context.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (Q8_0)
Phi-4 14B
Microsoft's 14B that reasons above its weight, especially at math. 16K context.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (Q4_K_M)
Phi-4-mini 3.8B
3.8B with 128K context — a capable little reasoner for light rigs.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (Q8_0)
Qwen2.5 0.5B · 1.5B · 3B · 7B · 14B · 32B · 72B
Alibaba's strong multilingual family across every size from 0.5B to 72B.
runs on a Apple M1 · 8GB (Q5_K_M)
Qwen2.5-Coder 1.5B · 7B · 14B · 32B
Best open coding model class. The 32B rivals proprietary coders; 7B/14B fit most rigs.
runs on a Apple M1 · 8GB (Q8_0)
Qwen3 1.7B · 4B · 8B · 14B · 32B · 30B-A3B (MoE)
Latest Qwen generation with hybrid thinking modes. Strong reasoning per parameter.
runs on a Apple M1 · 8GB (Q4_K_M)
SmolLM2 1.7B
Tiny but surprisingly capable. Runs on almost anything, even CPU.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (FP16)
StarCoder2 3B · 7B · 15B
BigCode's permissive code models trained on The Stack v2.
runs on a Apple M1 · 8GB (Q4_K_M)
TinyLlama 1.1B
1.1B — the classic 'will it run on a potato' model. Yes, it will.
runs on a Apple M1 · 8GB (FP16)
Yi 1.5 6B · 9B · 34B
01.AI's bilingual models, strong at reasoning and code for their size.
runs on a NVIDIA GeForce RTX 3060 · 12GB VRAM (Q8_0)

Want a model added? Open an issue — ingestion pulls measured GGUF sizes from Hugging Face, so verdicts stay honest from day one.