Twoody

Local LLM calculator

Which local LLM can your computer run — and how fast?

Pick your Mac or PC. See which open-weight models fit its memory, how fast they write, how long they take to read a document, and what they are good for.

Which Mac do I have? Apple menu › About This Mac lists the chip and the memory. On a PC: Task Manager › Performance › GPU.

The models for this computer

On M6 with 24 GB: 20 of 36 models run well.

What Twoody installs on this Mac

Qwen3 14B, in one click, with conversations of 16k tokens.

Runs comfortably · 10–19 tokens/s (faster than you read) · 20 pages in 37 s

Download for Apple Silicon Mac

Run well on this computer

gpt-oss-20b

OpenAI · Apache 2.0

Via LM Studio, Ollama or mlx-lm

12.7 of 15 GB

Runs comfortably

30–54 tokens/s

near instant

20 pages in 12 s

  • Chat and writing
  • Your documents
  • Code
  • Reasoning
  • Agents
Llama 3.1 8B

Meta · Llama 3.1 Community

Via LM Studio, Ollama or mlx-lm

7.3 of 15 GB

Runs comfortably

18–33 tokens/s

near instant

20 pages in 20 s

  • Chat and writing
  • Your documents
  • Translation
Qwen3 8B

Qwen (Alibaba) · Apache 2.0

One click in Twoody

7.7 of 15 GB

Runs comfortably

18–32 tokens/s

near instant

20 pages in 21 s

  • Chat and writing
  • Your documents
  • Translation
  • Reasoning
DeepSeek-R1-0528-Qwen3-8B

DeepSeek · MIT

Via LM Studio, Ollama or mlx-lm

7.7 of 15 GB

Runs comfortably

18–32 tokens/s

near instant

20 pages in 21 s

  • Chat and writing
  • Your documents
  • Translation
  • Reasoning
Ministral 3 8B

Mistral AI · Apache 2.0

Via LM Studio, Ollama or mlx-lm

7.7 of 15 GB

Runs comfortably

17–31 tokens/s

near instant

20 pages in 22 s

  • Chat and writing
  • Your documents
  • Translation
Qwen3.5 9B

Qwen (Alibaba) · Apache 2.0

Via LM Studio, Ollama or mlx-lm

6.6 of 15 GB

Runs comfortably

14–25 tokens/s

near instant

20 pages in 25 s

  • Chat and writing
  • Your documents
  • Translation
  • Reasoning
Gemma 4 12B

Google · Apache 2.0

Via LM Studio, Ollama or mlx-lm

8.1 of 15 GB

Runs comfortably

13–23 tokens/s

faster than you read

20 pages in 30 s

  • Chat and writing
  • Your documents
  • Translation
Ministral 3 14B

Mistral AI · Apache 2.0

Via LM Studio, Ollama or mlx-lm

11.1 of 15 GB

Runs comfortably

11–20 tokens/s

faster than you read

20 pages in 35 s

  • Chat and writing
  • Your documents
  • Translation
Phi-4 14B

Microsoft · MIT

Via LM Studio, Ollama or mlx-lm

12.3 of 15 GB

Runs comfortably

11–19 tokens/s

faster than you read

20 pages in 37 s

  • Chat and writing
  • Your documents
  • Reasoning
Qwen3 14B

Qwen (Alibaba) · Apache 2.0

One click in Twoody

11.8 of 15 GB

Runs comfortably

10–19 tokens/s

faster than you read

20 pages in 37 s

  • Chat and writing
  • Your documents
  • Translation
  • Reasoning
Llama 3.2 3B

Meta · Llama 3.2 Community

Via LM Studio, Ollama or mlx-lm

4.2 of 15 GB

Runs comfortably

40–72 tokens/s

feels instant

20 pages in 8 s

  • Chat and writing
  • Translation
Phi-4-mini

Microsoft · MIT

Via LM Studio, Ollama or mlx-lm

4.9 of 15 GB

Runs comfortably

33–59 tokens/s

near instant

20 pages in 10 s

  • Chat and writing
Qwen3 4B

Qwen (Alibaba) · Apache 2.0

One click in Twoody

5.2 of 15 GB

Runs comfortably

32–58 tokens/s

near instant

20 pages in 10 s

  • Chat and writing
  • Translation
Gemma 4 E4B

Google · Apache 2.0

Via LM Studio, Ollama or mlx-lm

5.9 of 15 GB

Runs comfortably

30–53 tokens/s

near instant

20 pages in 11 s

  • Chat and writing
  • Translation
Qwen3.5 4B

Qwen (Alibaba) · Apache 2.0

Via LM Studio, Ollama or mlx-lm

3.7 of 15 GB

Runs comfortably

27–48 tokens/s

near instant

20 pages in 12 s

  • Chat and writing
  • Translation
Qwen3 0.6B

Qwen (Alibaba) · Apache 2.0

Via LM Studio, Ollama or mlx-lm

2.6 of 15 GB

Runs comfortably

81–244 tokens/s

feels instant

20 pages in 2 s

Qwen3 1.7B

Qwen (Alibaba) · Apache 2.0

Via LM Studio, Ollama or mlx-lm

3.3 of 15 GB

Runs comfortably

45–136 tokens/s

feels instant

20 pages in 4 s

Gemma 4 E2B

Google · Apache 2.0

Via LM Studio, Ollama or mlx-lm

3.8 of 15 GB

Runs comfortably

54–96 tokens/s

feels instant

20 pages in 6 s

Qwen3.5 2B

Qwen (Alibaba) · Apache 2.0

Via LM Studio, Ollama or mlx-lm

1.9 of 15 GB

Runs comfortably

53–95 tokens/s

feels instant

20 pages in 5 s

MiniCPM5 2B

OpenBMB · Apache 2.0

Via LM Studio, Ollama or mlx-lm

2.7 of 15 GB

Runs comfortably

47–83 tokens/s

feels instant

20 pages in 6 s

Run, but much slower (8)
GLM-4.7-Flash

Z.ai · MIT

Via LM Studio, Ollama or mlx-lm

19.2 of 15 GB

Partly on the processor: much slower

14–26 tokens/s

near instant

20 pages in 24 s

  • Chat and writing
  • Your documents
  • Translation
  • Code
  • Reasoning
  • Agents
Qwen3 30B-A3B

Qwen (Alibaba) · Apache 2.0

Via LM Studio, Ollama or mlx-lm

20.1 of 15 GB

Partly on the processor: much slower

13–23 tokens/s

faster than you read

20 pages in 26 s

  • Chat and writing
  • Your documents
  • Translation
  • Code
  • Reasoning
  • Agents
Qwen3-Coder 30B-A3B

Qwen (Alibaba) · Apache 2.0

Via LM Studio, Ollama or mlx-lm

20.1 of 15 GB

Partly on the processor: much slower

13–23 tokens/s

faster than you read

20 pages in 26 s

  • Chat and writing
  • Your documents
  • Translation
  • Code
  • Agents
Gemma 4 26B-A4B

Google · Apache 2.0

Via LM Studio, Ollama or mlx-lm

17.7 of 15 GB

Partly on the processor: much slower

12–21 tokens/s

faster than you read

20 pages in 32 s

  • Chat and writing
  • Your documents
  • Translation
  • Code
  • Agents
Mistral Small 3.2 24B

Mistral AI · Apache 2.0

Via LM Studio, Ollama or mlx-lm

17 of 15 GB

Partly on the processor: much slower

2.7–4.8 tokens/s

slower than you read

20 pages in 3 min

  • Chat and writing
  • Your documents
  • Translation
  • Code
  • Agents
Devstral Small 2 24B

Mistral AI · Apache 2.0

Via LM Studio, Ollama or mlx-lm

17 of 15 GB

Partly on the processor: much slower

2.7–4.8 tokens/s

slower than you read

20 pages in 3 min

  • Chat and writing
  • Your documents
  • Translation
  • Code
  • Agents
Gemma 4 31B

Google · Apache 2.0

Via LM Studio, Ollama or mlx-lm

19.8 of 15 GB

Partly on the processor: much slower

2.1–3.8 tokens/s

slower than you read

20 pages in 3 min

  • Chat and writing
  • Your documents
  • Translation
  • Code
  • Agents
Qwen3.8 27B

Qwen (Alibaba) · Apache 2.0

Via LM Studio, Ollama or mlx-lm

17.6 of 15 GB

Partly on the processor: much slower

2–3.5 tokens/s

slower than you read

20 pages in 3 min

  • Chat and writing
  • Your documents
  • Translation
  • Code
  • Reasoning
  • Agents
Too big for this computer (8)
  • Qwen3.6 35B-A3B · 22.5 of 15 GB
  • Qwen3 32B · 23.8 of 15 GB
  • Llama 3.3 70B · 47.1 of 15 GB
  • Qwen3-Next 80B-A3B · 48.3 of 15 GB
  • gpt-oss-120b · 63.1 of 15 GB
  • Mistral Small 4 119B · 71.1 of 15 GB
  • Qwen3.5 122B-A10B · 75.7 of 15 GB
  • Qwen3 235B-A22B · 142.4 of 15 GB

Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.

How much memory do you need for a local LLM?

Memory decides which models fit: the model file, plus the conversation it keeps in mind. A Mac lets its graphics use about two thirds of its memory up to 32 GB, three quarters above. With 4-bit files, what runs comfortably:

MemoryRecommended — runs comfortably
8 GBLlama 3.2 3B, Phi-4-mini, Qwen3.5 4B
16 GBLlama 3.1 8B, Qwen3 8B, DeepSeek-R1-0528-Qwen3-8B
24 GBgpt-oss-20b, Llama 3.1 8B, Qwen3 8B
32 GBGLM-4.7-Flash, Qwen3 30B-A3B, Qwen3-Coder 30B-A3B
48 GBGLM-4.7-Flash, Qwen3 30B-A3B, Qwen3-Coder 30B-A3B
64 GBLlama 3.3 70B, GLM-4.7-Flash, Qwen3 30B-A3B
128 GBQwen3-Next 80B-A3B, gpt-oss-120b, Mistral Small 4 119B
256 GBQwen3 235B-A22B, gpt-oss-120b, Mistral Small 4 119B
512 GBQwen3 235B-A22B, gpt-oss-120b, Mistral Small 4 119B

Why memory bandwidth sets the speed

To write each word, the computer reads the model's active weights from memory. Speed therefore follows memory bandwidth more than processor power. Qwen3 8B, 4-bit, on a few machines:

Computer Memory bandwidth Writes
M1 68 GB/s 8.7–12 tokens/s
M2 100 GB/s 13–19 tokens/s
M4 120 GB/s 15–21 tokens/s
M5 153 GB/s 19–27 tokens/s
M4 Pro 273 GB/s 31–45 tokens/s
M4 Max 546 GB/s 53–76 tokens/s
M5 Max 614 GB/s 72–104 tokens/s
GeForce RTX 4090 (24 GB) 1,008 GB/s 100+ tokens/s

4-bit or 8-bit?

Models are shared compressed. 4-bit files (Q4_K_M, or MXFP4 for gpt-oss) are half the size of 8-bit ones for a small loss in quality: it is what Twoody installs, and the right choice on most computers. 8-bit is a little finer and needs twice the memory.

Tokens per second: what it feels like

A token is about three quarters of a word. You read silently at 5 to 6 tokens per second.

under 5slower than you read
5 to 10about your reading speed
10 to 20faster than you read
20 to 50near instant
over 50feels instant

A model that thinks first writes 300 to 3,000 tokens of reasoning before its answer: at 10 tokens per second, from half a minute to five minutes. In Twoody, the Fast, Thinking and Deep settings choose how long it thinks.

Reading a document

Before answering about a document, the model reads it — and that is where recent chips gain the most. Qwen3 8B reading a 20-page document, with 16 GB:

M13 min
M476 s
M527 s

Mac or PC?

A Mac shares its memory between the processor and the graphics: a 64 GB Mac runs models no 24 GB graphics card holds. A PC graphics card has less memory but reads it very fast: with a model that fits, an RTX 4090 writes faster than any Mac. Twoody runs on Macs today; the Windows and Linux versions are in preparation.

The best local LLM, by computer

Method and sources

Each machine's speed is computed from its measured speed on llama.cpp's reference model (LLaMA 2 7B, 4-bit), scaled by what each model reads per word, with the memory bandwidths of Apple's and NVIDIA's specifications. Checked against 20 published measurements that were not used to build it: median error 4.5 %, largest 23 %, all within the ranges shown.

Figures reviewed on September 26, 2026.

Sources

Download the data (JSON, CC BY 4.0)

Questions about running an LLM locally

Can I run a local LLM with 8 GB of memory?

Yes, a small one: Qwen3 4B or Qwen3.5 4B run on an 8 GB Mac and are enough for emails, summaries and simple questions. Close heavy apps while you use it. Twoody installs Qwen3 4B on these Macs.

Which LLM for a Mac with 16 GB?

A model of 7 to 9 billion parameters: Qwen3 8B, Qwen3.5 9B or Llama 3.1 8B write about 15 to 25 tokens per second on recent Macs and answer well about your documents. Twoody installs Qwen3 8B.

How much memory does a 70B model need?

About 48 GB for its 4-bit file and a short conversation: a Mac with 64 GB runs Llama 3.3 70B, at 7 to 13 tokens per second depending on the chip. Mixtures of experts such as gpt-oss-120b need more memory but write much faster.

Is a 4-bit model much worse than an 8-bit one?

Slightly. For writing, summaries and questions about documents the difference is hard to notice, and 4-bit halves the memory and doubles the speed. Choose 8-bit only when memory is plentiful.

Mac or NVIDIA graphics card for a local LLM?

A graphics card writes faster with a model that fits its memory (8 to 32 GB); a Mac holds much larger models in its unified memory (up to 512 GB). For one person working on documents, a Mac with 16 to 64 GB is the simplest choice.

Does a local LLM make a MacBook Air too hot?

A MacBook Air has no fan: after a few minutes of continuous writing it slows down to stay cool. For answers of normal length it is not noticeable; for long sessions, a MacBook Pro or a Mac mini keeps its speed.

Do I need LM Studio, Ollama or MLX?

Not with Twoody: it installs a model and runs it with its own engine, llama.cpp, in one click. If you already use LM Studio, Ollama or mlx-lm, Twoody uses their models too.

How accurate are these estimates?

On 20 published measurements that were not used to build the formula, the median error is 4.5 % and the largest 23 %; each fell within the range shown. Your speed also depends on the app, its version and what else the computer is doing.

Try local AI on your Mac, for free.

Version 0.12.1 · macOS 13 or later · no account, no subscription. Download it, install a model in one click, and ask your first question — even offline.