Local LLM calculator
The best local LLM for your iMac
From the M1 to the M4, with 8 to 32 GB: the strongest model that runs comfortably on each iMac, and how fast it writes.
| Your computer | Memory | Recommended — runs comfortably | Writes | One click in Twoody |
|---|---|---|---|---|
M4 — MacBook Air, MacBook Pro, Mac mini, iMac |
16 GB | Llama 3.1 8B | 15–22 tokens/s | Qwen3 8B |
| 24 GB | gpt-oss-20b | 25–36 tokens/s | Qwen3 14B | |
| 32 GB | GLM-4.7-Flash | 32–46 tokens/s | Qwen3 14B | |
M3 — MacBook Air, MacBook Pro, iMac |
8 GB | Llama 3.2 3B | 31–45 tokens/s | Qwen3 4B |
| 16 GB | Llama 3.1 8B | 13–19 tokens/s | Qwen3 8B | |
| 24 GB | gpt-oss-20b | 23–32 tokens/s | Qwen3 14B | |
M1 — MacBook Air, MacBook Pro, Mac mini, iMac |
8 GB | Llama 3.2 3B | 20–29 tokens/s | Qwen3 4B |
| 16 GB | Llama 3.1 8B | 8.9–13 tokens/s | Qwen3 8B |
Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.
An iMac with 8 GB runs models of about 4 billion parameters; with 16 or 24 GB, 8 to 14 billion — enough for writing and for questions about your documents.
The M3 and M4 write and read faster than the M1 with the same model; the memory decides which models fit.
On a desk all day, an iMac is also a good computer for the phone apps in preparation to use from afar.
Questions about running an LLM locally
Which LLM for an iMac with M4?
With 16 GB, Qwen3 8B; with 24 or 32 GB, Qwen3 14B or a mixture of experts such as gpt-oss-20b. Twoody installs the Qwen3 that suits its memory.
Can an iMac with 8 GB run a local LLM?
Yes, a model of about 4 billion parameters such as Qwen3 4B, enough for emails and summaries.
And an Intel iMac?
Twoody runs on Intel iMacs with macOS 13 or later; the model runs on the processor, more slowly. See the page on Intel Macs.
The best local LLM, by computer
Try local AI on your Mac, for free.
Version 0.11.5 · macOS 13 or later · no account, no subscription. Download it, install a model in one click, and ask your first question — even offline.