Twoody

Local LLM calculator

The best local LLM for your Mac Studio

From the M1 Max to the M5 Ultra with 512 GB: the largest models a Mac can hold, and how fast they write.

Your computer Memory Recommended — runs comfortably Writes One click in Twoody

M5 Max — MacBook Pro, Mac Studio

36 GB GLM-4.7-Flash 100+ tokens/s Qwen3 14B
48 GB GLM-4.7-Flash 100+ tokens/s Qwen3 14B
64 GB GLM-4.7-Flash 100+ tokens/s Qwen3 14B
128 GB Qwen3-Next 80B-A3B 96–193 tokens/s Qwen3 14B

M5 Ultra — Mac Studio

96 GB gpt-oss-120b 92–163 tokens/s Qwen3 14B
256 GB Qwen3 235B-A22B 28–50 tokens/s Qwen3 14B
512 GB Qwen3 235B-A22B 28–50 tokens/s Qwen3 14B

M3 Ultra — Mac Studio

96 GB gpt-oss-120b 63–90 tokens/s Qwen3 14B
256 GB Qwen3 235B-A22B 20–28 tokens/s Qwen3 14B
512 GB Qwen3 235B-A22B 20–28 tokens/s Qwen3 14B

M4 Max — MacBook Pro, Mac Studio

36 GB GLM-4.7-Flash 73–106 tokens/s Qwen3 14B
48 GB gpt-oss-20b 85–122 tokens/s Qwen3 14B
64 GB gpt-oss-20b 85–122 tokens/s Qwen3 14B
128 GB Qwen3-Next 80B-A3B 53–107 tokens/s Qwen3 14B

M2 Ultra — Mac Studio

64 GB gpt-oss-20b 89–128 tokens/s Qwen3 14B
128 GB gpt-oss-120b 63–91 tokens/s Qwen3 14B
192 GB Qwen3 235B-A22B 20–28 tokens/s Qwen3 14B

M2 Max — MacBook Pro, Mac Studio

32 GB gpt-oss-20b 62–90 tokens/s Qwen3 14B
64 GB gpt-oss-20b 62–90 tokens/s Qwen3 14B
96 GB Qwen3-Next 80B-A3B 39–78 tokens/s Qwen3 14B

M1 Ultra — Mac Studio

64 GB gpt-oss-20b 75–107 tokens/s Qwen3 14B
128 GB gpt-oss-120b 52–75 tokens/s Qwen3 14B

M1 Max — MacBook Pro, Mac Studio

32 GB gpt-oss-20b 56–80 tokens/s Qwen3 14B
64 GB gpt-oss-20b 56–80 tokens/s Qwen3 14B

Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.

With 128 to 512 GB of unified memory, a Mac Studio holds models no graphics card can: gpt-oss-120b, Qwen3.5 122B-A10B, even Qwen3 235B-A22B.

Its Max and Ultra chips read their memory at 400 GB/s to 1.2 TB/s: mixtures of experts answer at dozens of tokens per second.

Quiet even under load, it keeps its speed through long sessions.

Questions about running an LLM locally

Which LLM for a Mac Studio with an Ultra chip?

With 256 or 512 GB, Qwen3 235B-A22B, the largest model here, runs comfortably; gpt-oss-120b and Qwen3.5 122B-A10B write faster.

Is 128 GB enough for a 70B model?

Yes, comfortably: Llama 3.3 70B needs about 48 GB at 4-bit. With 128 GB, mixtures of experts such as gpt-oss-120b fit too.

Max or Ultra for a local LLM?

The Ultra doubles the memory bandwidth and the memory: faster answers with large models, and the largest models. For models up to 70B, a Max is enough.

Try local AI on your Mac, for free.

Version 0.11.5 · macOS 13 or later · no account, no subscription. Download it, install a model in one click, and ask your first question — even offline.