▲ runlocal
checker
models
matrix
new
graph
github
npm
Model catalog
37 curated open-weight models. Click any model to see what runs on your rig.
DeepSeek-R1 (distill)
1.5B · 7B · 8B (Llama) · 14B · 32B · 70B (Llama)
Open reasoning model. Distilled into Qwen/Llama backbones you can actually run locally.
reasoning
coding
general
Llama 3.1
8B · 70B
Meta's workhorse general model. Reliable all-rounder with 128K context.
chat
general
reasoning
agents
Qwen2.5
0.5B · 1.5B · 3B · 7B · 14B · 32B · 72B
Alibaba's strong multilingual family across every size from 0.5B to 72B.
chat
general
reasoning
agents
Qwen2.5-Coder
1.5B · 7B · 14B · 32B
Best open coding model class. The 32B rivals proprietary coders; 7B/14B fit most rigs.
coding
agents
Qwen3
1.7B · 4B · 8B · 14B · 32B · 30B-A3B (MoE)
Latest Qwen generation with hybrid thinking modes. Strong reasoning per parameter.
chat
general
reasoning
agents
coding
Gemma 4
12B
Google's Gemma 3 successor: unified any-to-any with vision, 256K context, Apache-2.0.
chat
general
vision
Llama 3.2
1B · 3B
Tiny Llamas for edge + on-device. Fast, light, 128K context.
chat
general
agents
Llama 3.3
70B
70B that rivals far larger models. The best open general model that fits a 64GB Mac.
chat
general
reasoning
agents
Mistral 7B
7B
The classic fast 7B. Apache-licensed, runs anywhere.
chat
general
agents
Mistral Small 3
24B
24B that punches at 70B level with much faster generation. 32K context.
chat
general
reasoning
agents
Gemma 3
1B · 4B · 12B · 27B
Google's latest, with vision and 128K context. Strong quality per parameter.
chat
general
vision
Mistral Nemo
12B
12B with 128K context, built with NVIDIA. Great mid-size all-rounder.
chat
general
agents
Gemma 2
2B · 9B · 27B
Google's efficient open models. The 9B is a standout for its size. 8K context.
chat
general
Phi-4
14B
Microsoft's 14B that reasons above its weight, especially at math. 16K context.
reasoning
general
coding
Nomic Embed Text
137M
Fast, high-quality text embeddings for RAG. Tiny — runs anywhere.
embeddings
rag
Mixtral 8x7B
8x7B (MoE)
Sparse MoE: 47B total but only ~13B active per token, so it decodes fast — if it fits in memory.
chat
general
reasoning
Codestral
22B
Mistral's 22B code model with strong fill-in-the-middle. 32K context.
coding
agents
Phi-4-mini
3.8B
3.8B with 128K context — a capable little reasoner for light rigs.
reasoning
general
chat
Llama 3.2 Vision
11B
Multimodal Llama — image understanding plus chat. 11B fits a 16GB+ Mac.
vision
chat
mxbai-embed-large
335M
Strong open embedding model for retrieval. Small and quick.
embeddings
rag
Phi-3.5-mini
3.8B
3.8B with 128K context. No GQA, so KV grows fast at long context.
general
chat
reasoning
SmolLM2
1.7B
Tiny but surprisingly capable. Runs on almost anything, even CPU.
chat
general
StarCoder2
3B · 7B · 15B
BigCode's permissive code models trained on The Stack v2.
coding
Code Llama
7B · 13B · 34B
Meta's Llama-2-based code models. Still popular on Ollama for completion.
coding
Yi 1.5
6B · 9B · 34B
01.AI's bilingual models, strong at reasoning and code for their size.
chat
general
reasoning
Command R
35B
Cohere's RAG/tool-use specialist. No GQA, so KV cache is heavy at long context.
rag
agents
general
LLaVA
7B · 13B
Popular open vision-language model for image Q&A and captioning.
vision
chat
TinyLlama
1.1B
1.1B — the classic 'will it run on a potato' model. Yes, it will.
chat
general
Granite 3.1 8B
8B
chat
coding
reasoning
rag
Falcon3 7B
7B
chat
coding
reasoning
OLMo 2 7B
7B
chat
general
Yi-Coder 9B
9B
coding
Nemotron-Mini 4B
4B
chat
reasoning
QwQ 32B
32B
reasoning
chat
Qwen3 Coder 30B-A3B
30B-A3B (MoE)
coding
agents
Granite 3.3 8B
8B
chat
rag
agents
SmolLM3 3B
3B
chat
general
reasoning