Local LLM calculator
The best local LLM for your MacBook Air
From the M1 to the M5, with 8 to 32 GB: the strongest model that runs comfortably on each MacBook Air, and how fast it writes.
| Your computer | Memory | Recommended — runs comfortably | Writes | One click in Twoody |
|---|---|---|---|---|
M5 — MacBook Air, MacBook Pro |
16 GB | Llama 3.1 8B | 20–28 tokens/s | Qwen3.5 9B |
| 24 GB | gpt-oss-20b | 32–47 tokens/s | Qwen3.5 9B | |
| 32 GB | GLM-4.7-Flash | 41–59 tokens/s | Qwen3.5 9B | |
M4 — MacBook Air, MacBook Pro, Mac mini, iMac |
16 GB | Llama 3.1 8B | 15–22 tokens/s | Qwen3.5 9B |
| 24 GB | gpt-oss-20b | 25–36 tokens/s | Qwen3.5 9B | |
| 32 GB | GLM-4.7-Flash | 32–46 tokens/s | Qwen3.5 9B | |
M3 — MacBook Air, MacBook Pro, iMac |
8 GB | Llama 3.2 3B | 31–45 tokens/s | Qwen3.5 4B |
| 16 GB | Llama 3.1 8B | 13–19 tokens/s | Qwen3.5 9B | |
| 24 GB | gpt-oss-20b | 23–32 tokens/s | Qwen3.5 9B | |
M2 — MacBook Air, MacBook Pro, Mac mini |
8 GB | Llama 3.2 3B | 32–45 tokens/s | Qwen3.5 4B |
| 16 GB | Llama 3.1 8B | 13–19 tokens/s | Qwen3.5 9B | |
| 24 GB | gpt-oss-20b | 23–33 tokens/s | Qwen3.5 9B | |
M1 — MacBook Air, MacBook Pro, Mac mini, iMac |
8 GB | Llama 3.2 3B | 20–29 tokens/s | Qwen3.5 4B |
| 16 GB | Llama 3.1 8B | 8.9–13 tokens/s | Qwen3.5 9B |
Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.
A MacBook Air shares its memory between the processor and the graphics. With 16 GB, a model of 8 billion parameters leaves room for your other apps; with 8 GB, stay with about 4 billion.
Speed follows the chip's memory bandwidth: each generation writes faster, and the M5 reads a document several times faster than the chips before it.
Without a fan, a MacBook Air slows down after a few minutes of continuous work: fine for answers of normal length, less so for long sessions.
Questions about running an LLM locally
Which LLM for a MacBook Air M4 with 16 GB?
Qwen3.5 9B or Qwen3 8B: they run comfortably and write about 15 to 21 tokens per second. Twoody installs Qwen3.5 9B in one click.
Can a MacBook Air with 8 GB run a local LLM?
Yes, a model of about 4 billion parameters, like Qwen3.5 4B, which Twoody installs on these Macs. Close heavy apps while using it.
Does a MacBook Air get too hot running a local LLM?
It slows down rather than overheats: without a fan, it lowers its speed after a few minutes of continuous writing. For everyday questions it is not noticeable.
The best local LLM, by computer
Twoody is in private beta.
On Mac, Windows and Linux, free, with no account and no subscription — and on iPhone and Android, with Twoody on your computer. Leave your email: we will write to you when Twoody opens to you.