Local LLM calculator
Which local LLM can your computer run — and how fast?
Pick your Mac or PC. See which open-weight models fit its memory, how fast they write, how long they take to read a document, and what they are good for.
The models for this computer
On a PC with GeForce RTX 3090 (24 GB) and 32 GB of RAM: 30 of 36 models run well.
Twoody on a PC
Twoody for Windows and Linux is in preparation; the version being tested runs the model on the processor. The speeds below are those of LM Studio or Ollama, which use the graphics card.
Run well on this computer
12.5 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 2 s
- Chat and writing
- Your documents
- Code
- Reasoning
- Agents
18.8 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 2 s
- Chat and writing
- Your documents
- Translation
- Code
- Reasoning
- Agents
22.3 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
- Chat and writing
- Your documents
- Translation
- Code
- Reasoning
- Agents
17.7 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 2 s
- Chat and writing
- Your documents
- Translation
- Code
- Agents
19.4 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 2 s
- Chat and writing
- Your documents
- Translation
- Code
- Reasoning
- Agents
19.4 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 2 s
- Chat and writing
- Your documents
- Translation
- Code
- Agents
15.8 of 22.6 GB
Runs comfortably
40–58 tokens/s
feels instant
12 pages in 6 s
- Chat and writing
- Your documents
- Translation
- Code
- Agents
15.8 of 22.6 GB
Runs comfortably
40–58 tokens/s
feels instant
12 pages in 6 s
- Chat and writing
- Your documents
- Translation
- Code
- Agents
17.1 of 22.6 GB
Runs comfortably
26–53 tokens/s
near instant
12 pages in 7 s
- Chat and writing
- Your documents
- Translation
- Code
- Reasoning
- Agents
19.5 of 22.6 GB
Runs comfortably
31–45 tokens/s
near instant
12 pages in 7 s
- Chat and writing
- Your documents
- Translation
- Code
- Agents
21.8 of 22.6 GB
Runs comfortably
29–41 tokens/s
near instant
12 pages in 8 s
- Chat and writing
- Your documents
- Translation
- Code
- Reasoning
- Agents
6.3 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 2 s
- Chat and writing
- Your documents
- Translation
6.5 of 22.6 GB
Runs comfortably
99–143 tokens/s
feels instant
12 pages in 2 s
- Chat and writing
- Your documents
- Translation
- Reasoning
6.5 of 22.6 GB
Runs comfortably
99–143 tokens/s
feels instant
12 pages in 2 s
- Chat and writing
- Your documents
- Translation
- Reasoning
6.6 of 22.6 GB
Runs comfortably
98–141 tokens/s
feels instant
12 pages in 2 s
- Chat and writing
- Your documents
- Translation
6.3 of 22.6 GB
Runs comfortably
72–144 tokens/s
feels instant
12 pages in 2 s
- Chat and writing
- Your documents
- Translation
- Reasoning
8 of 22.6 GB
Runs comfortably
73–104 tokens/s
feels instant
12 pages in 3 s
- Chat and writing
- Your documents
- Translation
9.8 of 22.6 GB
Runs comfortably
65–94 tokens/s
feels instant
12 pages in 3 s
- Chat and writing
- Your documents
- Translation
10.8 of 22.6 GB
Runs comfortably
61–87 tokens/s
feels instant
12 pages in 3 s
- Chat and writing
- Your documents
- Reasoning
10.6 of 22.6 GB
Runs comfortably
60–87 tokens/s
feels instant
12 pages in 4 s
- Chat and writing
- Your documents
- Translation
- Reasoning
3.4 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
- Chat and writing
- Translation
3.9 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
- Chat and writing
4.1 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
- Chat and writing
- Translation
3.4 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
- Chat and writing
- Translation
5.8 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
- Chat and writing
- Translation
1.8 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
2.5 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
1.8 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
3.8 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
2.4 of 22.6 GB
Runs comfortably
100+ tokens/s
feels instant
12 pages in 1 s
Run, but much slower (2)
48.1 of 22.6 GB
Partly on the processor: much slower
15–30 tokens/s
near instant
12 pages in 4 s
- Chat and writing
- Your documents
- Translation
- Code
- Reasoning
- Agents
44.6 of 22.6 GB
Partly on the processor: much slower
1.2–1.7 tokens/s
slower than you read
12 pages in 42 s
- Chat and writing
- Your documents
- Translation
- Code
- Reasoning
- Agents
Too big for this computer (4)
- gpt-oss-120b · 62.8 of 22.6 GB
- Mistral Small 4 119B · 70.9 of 22.6 GB
- Qwen3.5 122B-A10B · 75.5 of 22.6 GB
- Qwen3 235B-A22B · 141 of 22.6 GB
Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.
How much memory do you need for a local LLM?
Memory decides which models fit: the model file, plus the conversation it keeps in mind. A Mac lets its graphics use about two thirds of its memory up to 32 GB, three quarters above. With 4-bit files, what runs comfortably:
| Memory | Recommended — runs comfortably |
|---|---|
| 8 GB | Llama 3.2 3B, Phi-4-mini, Qwen3.5 4B |
| 16 GB | Llama 3.1 8B, Qwen3 8B, DeepSeek-R1-0528-Qwen3-8B |
| 24 GB | gpt-oss-20b, Llama 3.1 8B, Qwen3 8B |
| 32 GB | GLM-4.7-Flash, Qwen3 30B-A3B, Qwen3-Coder 30B-A3B |
| 48 GB | GLM-4.7-Flash, Qwen3 30B-A3B, Qwen3-Coder 30B-A3B |
| 64 GB | Llama 3.3 70B, GLM-4.7-Flash, Qwen3 30B-A3B |
| 128 GB | Qwen3-Next 80B-A3B, gpt-oss-120b, Mistral Small 4 119B |
| 256 GB | Qwen3 235B-A22B, gpt-oss-120b, Mistral Small 4 119B |
| 512 GB | Qwen3 235B-A22B, gpt-oss-120b, Mistral Small 4 119B |
Why memory bandwidth sets the speed
To write each word, the computer reads the model's active weights from memory. Speed therefore follows memory bandwidth more than processor power. Qwen3 8B, 4-bit, on a few machines:
| Computer | Memory bandwidth | Writes |
|---|---|---|
| M1 | 68 GB/s | 8.7–12 tokens/s |
| M2 | 100 GB/s | 13–19 tokens/s |
| M4 | 120 GB/s | 15–21 tokens/s |
| M5 | 153 GB/s | 19–27 tokens/s |
| M4 Pro | 273 GB/s | 31–45 tokens/s |
| M4 Max | 546 GB/s | 53–76 tokens/s |
| M5 Max | 614 GB/s | 72–104 tokens/s |
| GeForce RTX 4090 (24 GB) | 1,008 GB/s | 100+ tokens/s |
4-bit or 8-bit?
Models are shared compressed. 4-bit files (Q4_K_M, or MXFP4 for gpt-oss) are half the size of 8-bit ones for a small loss in quality: it is what Twoody installs, and the right choice on most computers. 8-bit is a little finer and needs twice the memory.
Tokens per second: what it feels like
A token is about three quarters of a word. You read silently at 5 to 6 tokens per second.
| under 5 | slower than you read |
| 5 to 10 | about your reading speed |
| 10 to 20 | faster than you read |
| 20 to 50 | near instant |
| over 50 | feels instant |
A model that thinks first writes 300 to 3,000 tokens of reasoning before its answer: at 10 tokens per second, from half a minute to five minutes. In Twoody, the Fast, Thinking and Deep settings choose how long it thinks.
Reading a document
Before answering about a document, the model reads it — and that is where recent chips gain the most. Qwen3 8B reading a 20-page document, with 16 GB:
| M1 | 3 min |
| M4 | 76 s |
| M5 | 27 s |
Mac or PC?
A Mac shares its memory between the processor and the graphics: a 64 GB Mac runs models no 24 GB graphics card holds. A PC graphics card has less memory but reads it very fast: with a model that fits, an RTX 4090 writes faster than any Mac. Twoody runs on Macs today; the Windows and Linux versions are in preparation.
The best local LLM, by computer
Method and sources
Each machine's speed is computed from its measured speed on llama.cpp's reference model (LLaMA 2 7B, 4-bit), scaled by what each model reads per word, with the memory bandwidths of Apple's and NVIDIA's specifications. Checked against 20 published measurements that were not used to build it: median error 4.5 %, largest 23 %, all within the ranges shown.
Figures reviewed on September 26, 2026.
Sources
- www.apple.com/newsroom/
- support.apple.com/specs
- github.com/ggml-org/llama.cpp/discussions/4167
- github.com/ggml-org/llama.cpp/discussions/15013
- github.com/ggml-org/llama.cpp/discussions/10879
- github.com/ggml-org/llama.cpp/discussions/15396
- www.nvidia.com/en-us/geforce/graphics-cards/compare/
- huggingface.co/models
- www.sciencedirect.com/science/article/abs/pii/S0749596X19300786
Questions about running an LLM locally
Can I run a local LLM with 8 GB of memory?
Yes, a small one: Qwen3 4B or Qwen3.5 4B run on an 8 GB Mac and are enough for emails, summaries and simple questions. Close heavy apps while you use it. Twoody installs Qwen3 4B on these Macs.
Which LLM for a Mac with 16 GB?
A model of 7 to 9 billion parameters: Qwen3 8B, Qwen3.5 9B or Llama 3.1 8B write about 15 to 25 tokens per second on recent Macs and answer well about your documents. Twoody installs Qwen3 8B.
How much memory does a 70B model need?
About 48 GB for its 4-bit file and a short conversation: a Mac with 64 GB runs Llama 3.3 70B, at 7 to 13 tokens per second depending on the chip. Mixtures of experts such as gpt-oss-120b need more memory but write much faster.
Is a 4-bit model much worse than an 8-bit one?
Slightly. For writing, summaries and questions about documents the difference is hard to notice, and 4-bit halves the memory and doubles the speed. Choose 8-bit only when memory is plentiful.
Mac or NVIDIA graphics card for a local LLM?
A graphics card writes faster with a model that fits its memory (8 to 32 GB); a Mac holds much larger models in its unified memory (up to 512 GB). For one person working on documents, a Mac with 16 to 64 GB is the simplest choice.
Does a local LLM make a MacBook Air too hot?
A MacBook Air has no fan: after a few minutes of continuous writing it slows down to stay cool. For answers of normal length it is not noticeable; for long sessions, a MacBook Pro or a Mac mini keeps its speed.
Do I need LM Studio, Ollama or MLX?
Not with Twoody: it installs a model and runs it with its own engine, llama.cpp, in one click. If you already use LM Studio, Ollama or mlx-lm, Twoody uses their models too.
How accurate are these estimates?
On 20 published measurements that were not used to build the formula, the median error is 4.5 % and the largest 23 %; each fell within the range shown. Your speed also depends on the app, its version and what else the computer is doing.
Try local AI on your Mac, for free.
Version 0.12.1 · macOS 13 or later · no account, no subscription. Download it, install a model in one click, and ask your first question — even offline.