Twoody

Local LLM calculator

The best local LLM for your Mac Studio

From the M1 Max to the M5 Ultra with 512 GB: the largest models a Mac can hold, and how fast they write.

Your computer Memory Recommended — runs comfortably Writes One click in Twoody

M5 Max — MacBook Pro, Mac Studio

36 GB GLM-4.7-Flash 100+ tokens/s Qwen3.5 9B
48 GB GLM-4.7-Flash 100+ tokens/s Qwen3.5 9B
64 GB GLM-4.7-Flash 100+ tokens/s Qwen3.5 9B
128 GB Qwen3-Next 80B-A3B 96–193 tokens/s Qwen3.5 9B

M5 Ultra — Mac Studio

96 GB gpt-oss-120b 92–163 tokens/s Qwen3.5 9B
256 GB Qwen3 235B-A22B 28–50 tokens/s Qwen3.5 9B
512 GB Qwen3 235B-A22B 28–50 tokens/s Qwen3.5 9B

M3 Ultra — Mac Studio

96 GB gpt-oss-120b 63–90 tokens/s Qwen3.5 9B
256 GB Qwen3 235B-A22B 20–28 tokens/s Qwen3.5 9B
512 GB Qwen3 235B-A22B 20–28 tokens/s Qwen3.5 9B

M4 Max — MacBook Pro, Mac Studio

36 GB GLM-4.7-Flash 73–106 tokens/s Qwen3.5 9B
48 GB gpt-oss-20b 85–122 tokens/s Qwen3.5 9B
64 GB gpt-oss-20b 85–122 tokens/s Qwen3.5 9B
128 GB Qwen3-Next 80B-A3B 53–107 tokens/s Qwen3.5 9B

M2 Ultra — Mac Studio

64 GB gpt-oss-20b 89–128 tokens/s Qwen3.5 9B
128 GB gpt-oss-120b 63–91 tokens/s Qwen3.5 9B
192 GB Qwen3 235B-A22B 20–28 tokens/s Qwen3.5 9B

M2 Max — MacBook Pro, Mac Studio

32 GB gpt-oss-20b 62–90 tokens/s Qwen3.5 9B
64 GB gpt-oss-20b 62–90 tokens/s Qwen3.5 9B
96 GB Qwen3-Next 80B-A3B 39–78 tokens/s Qwen3.5 9B

M1 Ultra — Mac Studio

64 GB gpt-oss-20b 75–107 tokens/s Qwen3.5 9B
128 GB gpt-oss-120b 52–75 tokens/s Qwen3.5 9B

M1 Max — MacBook Pro, Mac Studio

32 GB gpt-oss-20b 56–80 tokens/s Qwen3.5 9B
64 GB gpt-oss-20b 56–80 tokens/s Qwen3.5 9B

Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.

With 128 to 512 GB of unified memory, a Mac Studio holds models no graphics card can: gpt-oss-120b, Qwen3.5 122B-A10B, even Qwen3 235B-A22B.

Its Max and Ultra chips read their memory at 400 GB/s to 1.2 TB/s: mixtures of experts answer at dozens of tokens per second.

Quiet even under load, it keeps its speed through long sessions.

Questions about running an LLM locally

Which LLM for a Mac Studio with an Ultra chip?

With 256 or 512 GB, Qwen3 235B-A22B, the largest model here, runs comfortably; gpt-oss-120b and Qwen3.5 122B-A10B write faster.

Is 128 GB enough for a 70B model?

Yes, comfortably: Llama 3.3 70B needs about 48 GB at 4-bit. With 128 GB, mixtures of experts such as gpt-oss-120b fit too.

Max or Ultra for a local LLM?

The Ultra doubles the memory bandwidth and the memory: faster answers with large models, and the largest models. For models up to 70B, a Max is enough.

Twoody is in private beta.

On Mac, Windows and Linux, free, with no account and no subscription — and on iPhone and Android, with Twoody on your computer. Leave your email: we will write to you when Twoody opens to you.