Local LLM calculator
The best local LLM for your MacBook Air
From the M1 to the M5, with 8 to 32 GB: the strongest model that runs comfortably on each MacBook Air, and how fast it writes.
| Your computer | Memory | Recommended — runs comfortably | Writes | One click in Twoody |
|---|---|---|---|---|
M5 — MacBook Air, MacBook Pro |
16 GB | Llama 3.1 8B | 20–28 tokens/s | Qwen3.5 9B |
| 24 GB | gpt-oss-20b | 32–47 tokens/s | Qwen3.5 9B | |
| 32 GB | GLM-4.7-Flash | 41–59 tokens/s | Qwen3.5 9B | |
M4 — MacBook Air, MacBook Pro, Mac mini, iMac |
16 GB | Llama 3.1 8B | 15–22 tokens/s | Qwen3.5 9B |
| 24 GB | gpt-oss-20b | 25–36 tokens/s | Qwen3.5 9B | |
| 32 GB | GLM-4.7-Flash | 32–46 tokens/s | Qwen3.5 9B | |
M3 — MacBook Air, MacBook Pro, iMac |
8 GB | Llama 3.2 3B | 31–45 tokens/s | Qwen3.5 4B |
| 16 GB | Llama 3.1 8B | 13–19 tokens/s | Qwen3.5 9B | |
| 24 GB | gpt-oss-20b | 23–32 tokens/s | Qwen3.5 9B | |
M2 — MacBook Air, MacBook Pro, Mac mini |
8 GB | Llama 3.2 3B | 32–45 tokens/s | Qwen3.5 4B |
| 16 GB | Llama 3.1 8B | 13–19 tokens/s | Qwen3.5 9B | |
| 24 GB | gpt-oss-20b | 23–33 tokens/s | Qwen3.5 9B | |
M1 — MacBook Air, MacBook Pro, Mac mini, iMac |
8 GB | Llama 3.2 3B | 20–29 tokens/s | Qwen3.5 4B |
| 16 GB | Llama 3.1 8B | 8.9–13 tokens/s | Qwen3.5 9B |
Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.
A MacBook Air shares its memory between the processor and the graphics. With 16 GB, a model of 8 billion parameters leaves room for your other apps; with 8 GB, stay with about 4 billion.
Speed follows the chip's memory bandwidth: each generation writes faster, and the M5 reads a document several times faster than the chips before it.
Without a fan, a MacBook Air slows down after a few minutes of continuous work: fine for answers of normal length, less so for long sessions.
Questions about running an LLM locally
Which LLM for a MacBook Air M4 with 16 GB?
Qwen3.5 9B or Qwen3 8B: they run comfortably and write about 15 to 21 tokens per second. Twoody installs Qwen3.5 9B in one click.
Can a MacBook Air with 8 GB run a local LLM?
Yes, a model of about 4 billion parameters, like Qwen3.5 4B, which Twoody installs on these Macs. Close heavy apps while using it.
Does a MacBook Air get too hot running a local LLM?
It slows down rather than overheats: without a fan, it lowers its speed after a few minutes of continuous writing. For everyday questions it is not noticeable.
The best local LLM, by computer
Try local AI on your Mac, for free.
Version 0.13.1 · macOS 13 or later · no account, no subscription. Download it, install a model in one click, and ask your first question — even offline.