Twoody

Local LLM calculator

The best local LLM for your Mac mini

From the M1 to the M6, and the Pro versions: the strongest model that runs comfortably on each Mac mini, and how fast it writes.

Your computer Memory Recommended — runs comfortably Writes One click in Twoody

M5 Pro — MacBook Pro, Mac mini

24 GB gpt-oss-20b 68–97 tokens/s Qwen3 14B
48 GB GLM-4.7-Flash 92–132 tokens/s Qwen3 14B
64 GB GLM-4.7-Flash 92–132 tokens/s Qwen3 14B

M6 — Mac mini

16 GB Llama 3.1 8B 17–31 tokens/s Qwen3 8B
24 GB gpt-oss-20b 30–54 tokens/s Qwen3 14B
32 GB GLM-4.7-Flash 36–64 tokens/s Qwen3 14B

M4 Pro — MacBook Pro, Mac mini

24 GB gpt-oss-20b 52–74 tokens/s Qwen3 14B
48 GB GLM-4.7-Flash 56–81 tokens/s Qwen3 14B
64 GB GLM-4.7-Flash 56–81 tokens/s Qwen3 14B

M4 — MacBook Air, MacBook Pro, Mac mini, iMac

16 GB Llama 3.1 8B 15–22 tokens/s Qwen3 8B
24 GB gpt-oss-20b 25–36 tokens/s Qwen3 14B
32 GB GLM-4.7-Flash 32–46 tokens/s Qwen3 14B

M2 Pro — MacBook Pro, Mac mini

16 GB Llama 3.1 8B 24–35 tokens/s Qwen3 8B
32 GB GLM-4.7-Flash 45–65 tokens/s Qwen3 14B

M2 — MacBook Air, MacBook Pro, Mac mini

8 GB Llama 3.2 3B 32–45 tokens/s Qwen3 4B
16 GB Llama 3.1 8B 13–19 tokens/s Qwen3 8B
24 GB gpt-oss-20b 23–33 tokens/s Qwen3 14B

M1 — MacBook Air, MacBook Pro, Mac mini, iMac

8 GB Llama 3.2 3B 20–29 tokens/s Qwen3 4B
16 GB Llama 3.1 8B 8.9–13 tokens/s Qwen3 8B

Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.

The Mac mini is the least expensive way to run a local LLM on a Mac: a model of 8 billion parameters with 16 GB, a mixture of experts with 48 GB and a Pro chip.

Silent and always on, a Mac mini is also the kind of computer the phone apps in preparation are made to use from afar.

The Pro versions read their memory more than twice as fast as the base chips: larger models, and faster answers.

Questions about running an LLM locally

Which LLM for a Mac mini M4 with 16 GB?

Qwen3 8B, Qwen3.5 9B or Llama 3.1 8B run comfortably, at about 15 to 21 tokens per second. Twoody installs Qwen3 8B in one click.

Is a Mac mini M4 Pro worth it for local AI?

With 48 or 64 GB, yes: it runs mixtures of experts such as Qwen3 30B-A3B or gpt-oss-20b at several dozen tokens per second.

Can a Mac mini serve a local LLM to other devices?

That is what the phone apps in preparation will do: use the model of your Mac from an iPhone or an Android phone, with an end-to-end encrypted connection.

Try local AI on your Mac, for free.

Version 0.11.5 · macOS 13 or later · no account, no subscription. Download it, install a model in one click, and ask your first question — even offline.