Local LLM calculator
The best local LLM for your Mac mini
From the M1 to the M6, and the Pro versions: the strongest model that runs comfortably on each Mac mini, and how fast it writes.
| Your computer | Memory | Recommended — runs comfortably | Writes | One click in Twoody |
|---|---|---|---|---|
M5 Pro — MacBook Pro, Mac mini |
24 GB | gpt-oss-20b | 68–97 tokens/s | Qwen3.5 9B |
| 48 GB | GLM-4.7-Flash | 92–132 tokens/s | Qwen3.5 9B | |
| 64 GB | GLM-4.7-Flash | 92–132 tokens/s | Qwen3.5 9B | |
M6 — Mac mini |
16 GB | Llama 3.1 8B | 17–31 tokens/s | Qwen3.5 9B |
| 24 GB | gpt-oss-20b | 30–54 tokens/s | Qwen3.5 9B | |
| 32 GB | GLM-4.7-Flash | 36–64 tokens/s | Qwen3.5 9B | |
M4 Pro — MacBook Pro, Mac mini |
24 GB | gpt-oss-20b | 52–74 tokens/s | Qwen3.5 9B |
| 48 GB | GLM-4.7-Flash | 56–81 tokens/s | Qwen3.5 9B | |
| 64 GB | GLM-4.7-Flash | 56–81 tokens/s | Qwen3.5 9B | |
M4 — MacBook Air, MacBook Pro, Mac mini, iMac |
16 GB | Llama 3.1 8B | 15–22 tokens/s | Qwen3.5 9B |
| 24 GB | gpt-oss-20b | 25–36 tokens/s | Qwen3.5 9B | |
| 32 GB | GLM-4.7-Flash | 32–46 tokens/s | Qwen3.5 9B | |
M2 Pro — MacBook Pro, Mac mini |
16 GB | Llama 3.1 8B | 24–35 tokens/s | Qwen3.5 9B |
| 32 GB | GLM-4.7-Flash | 45–65 tokens/s | Qwen3.5 9B | |
M2 — MacBook Air, MacBook Pro, Mac mini |
8 GB | Llama 3.2 3B | 32–45 tokens/s | Qwen3.5 4B |
| 16 GB | Llama 3.1 8B | 13–19 tokens/s | Qwen3.5 9B | |
| 24 GB | gpt-oss-20b | 23–33 tokens/s | Qwen3.5 9B | |
M1 — MacBook Air, MacBook Pro, Mac mini, iMac |
8 GB | Llama 3.2 3B | 20–29 tokens/s | Qwen3.5 4B |
| 16 GB | Llama 3.1 8B | 8.9–13 tokens/s | Qwen3.5 9B |
Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.
The Mac mini is the least expensive way to run a local LLM on a Mac: a model of 8 billion parameters with 16 GB, a mixture of experts with 48 GB and a Pro chip.
Silent and always on, a Mac mini is also the kind of computer the iPhone and Android apps, in private beta, are made to use from afar.
The Pro versions read their memory more than twice as fast as the base chips: larger models, and faster answers.
Questions about running an LLM locally
Which LLM for a Mac mini M4 with 16 GB?
Qwen3.5 9B, Qwen3 8B or Llama 3.1 8B run comfortably, at about 15 to 21 tokens per second. Twoody installs Qwen3.5 9B in one click.
Is a Mac mini M4 Pro worth it for local AI?
With 48 or 64 GB, yes: it runs mixtures of experts such as Qwen3 30B-A3B or gpt-oss-20b at several dozen tokens per second.
Can a Mac mini serve a local LLM to other devices?
That is what the iPhone and Android apps, in private beta, do: they use the model of your Mac from the phone.
The best local LLM, by computer
Twoody is in private beta.
On Mac, Windows and Linux, free, with no account and no subscription — and on iPhone and Android, with Twoody on your computer. Leave your email: we will write to you when Twoody opens to you.