Local LLM calculator
The best local LLM for your NVIDIA graphics card
From the RTX 3060 to the RTX 5090: the strongest model that fits each card's memory, and how fast it writes.
| Your computer | Memory | Recommended — runs comfortably | Writes |
|---|---|---|---|
GeForce RTX 5090 (32 GB) |
32 GB | gpt-oss-20b | 100+ tokens/s |
GeForce RTX 5080 (16 GB) |
16 GB | gpt-oss-20b | 100+ tokens/s |
GeForce RTX 5070 Ti (16 GB) |
16 GB | gpt-oss-20b | 100+ tokens/s |
GeForce RTX 5070 (12 GB) |
12 GB | Llama 3.1 8B | 80–116 tokens/s |
GeForce RTX 5060 Ti (16 GB) |
16 GB | gpt-oss-20b | 82–142 tokens/s |
GeForce RTX 5060 Ti (8 GB) |
8 GB | Llama 3.1 8B | 57–83 tokens/s |
GeForce RTX 5060 (8 GB) |
8 GB | Llama 3.1 8B | 50–89 tokens/s |
GeForce RTX 4090 (24 GB) |
24 GB | GLM-4.7-Flash | 100+ tokens/s |
GeForce RTX 4080 SUPER (16 GB) |
16 GB | gpt-oss-20b | 100+ tokens/s |
GeForce RTX 4080 (16 GB) |
16 GB | gpt-oss-20b | 100+ tokens/s |
GeForce RTX 4070 Ti SUPER (16 GB) |
16 GB | gpt-oss-20b | 100+ tokens/s |
GeForce RTX 4070 Ti (12 GB) |
12 GB | Llama 3.1 8B | 66–95 tokens/s |
GeForce RTX 4070 SUPER (12 GB) |
12 GB | Llama 3.1 8B | 66–95 tokens/s |
GeForce RTX 4070 (12 GB) |
12 GB | Llama 3.1 8B | 58–84 tokens/s |
GeForce RTX 4060 Ti (16 GB) |
16 GB | gpt-oss-20b | 55–96 tokens/s |
GeForce RTX 4060 Ti (8 GB) |
8 GB | Llama 3.1 8B | 38–55 tokens/s |
GeForce RTX 4060 (8 GB) |
8 GB | Llama 3.1 8B | 31–56 tokens/s |
GeForce RTX 3090 (24 GB) |
24 GB | gpt-oss-20b | 100+ tokens/s |
GeForce RTX 3080 (10 GB) |
10 GB | Llama 3.1 8B | 88–127 tokens/s |
GeForce RTX 3070 (8 GB) |
8 GB | Llama 3.1 8B | 50–72 tokens/s |
GeForce RTX 3060 Ti (8 GB) |
8 GB | Llama 3.1 8B | 54–78 tokens/s |
GeForce RTX 3060 (12 GB) |
12 GB | Llama 3.1 8B | 47–68 tokens/s |
Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.
On a PC, the graphics card's memory (VRAM) decides: 8 GB for models of 4 to 8 billion parameters, 16 GB for 14 billion or gpt-oss-20b, 24 to 32 GB for 32 billion.
Graphics cards read their memory very fast: with a model that fits, they write faster than any Mac.
Twoody for Windows and Linux is in private beta: on Windows 10 and 11 (x64), it runs the model on an NVIDIA card through Vulkan when the driver supports it, and on the processor otherwise; on Linux, on the processor. The speeds on this page are those of LM Studio or Ollama, not measured with Twoody.
Questions about running an LLM locally
Which LLM for an RTX 4090?
With 24 GB, Qwen3 32B or Gemma 4 31B fit at 4-bit, and gpt-oss-20b writes at well over 100 tokens per second.
Is an 8 GB graphics card enough?
For models of 4 to 8 billion parameters, yes, with a short conversation. Larger ones spill into the computer's memory and slow down a lot.
Does Twoody run on Windows with an NVIDIA card?
In the private beta, yes: on Windows 10 and 11 (x64), Twoody runs the model on an NVIDIA card through Vulkan when the driver supports it, and on the processor otherwise. Leave your email on the platforms page to join the waitlist.
The best local LLM, by computer
Twoody is in private beta.
On Mac, Windows and Linux, free, with no account and no subscription — and on iPhone and Android, with Twoody on your computer. Leave your email: we will write to you when Twoody opens to you.