Local LLM calculator
The best local LLM for your AMD graphics card
From the RX 7600 to the RX 9070 XT: the strongest model that fits each card's memory, and how fast it writes.
| Your computer | Memory | Recommended — runs comfortably | Writes |
|---|---|---|---|
Radeon RX 9070 XT (16 GB) |
16 GB | gpt-oss-20b | 100+ tokens/s |
Radeon RX 9070 (16 GB) |
16 GB | gpt-oss-20b | 100+ tokens/s |
Radeon RX 9060 XT (16 GB) |
16 GB | gpt-oss-20b | 61–107 tokens/s |
Radeon RX 9060 XT (8 GB) |
8 GB | Llama 3.1 8B | 42–61 tokens/s |
Radeon RX 7900 XTX (24 GB) |
24 GB | GLM-4.7-Flash | 100+ tokens/s |
Radeon RX 7900 XT (20 GB) |
20 GB | gpt-oss-20b | 100+ tokens/s |
Radeon RX 7900 GRE (16 GB) |
16 GB | gpt-oss-20b | 100+ tokens/s |
Radeon RX 7800 XT (16 GB) |
16 GB | gpt-oss-20b | 100+ tokens/s |
Radeon RX 7600 XT (16 GB) |
16 GB | gpt-oss-20b | 46–80 tokens/s |
Radeon RX 7600 (8 GB) |
8 GB | Llama 3.1 8B | 36–52 tokens/s |
Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.
The card's memory decides which models fit: 16 GB runs models of 14 billion parameters or gpt-oss-20b; the 24 GB RX 7900 XTX, 32 billion.
llama.cpp runs Radeon cards through Vulkan: writing is fast; reading long documents is slower than on NVIDIA cards.
Twoody for Windows and Linux is in preparation; the version being tested runs models on the processor. Meanwhile, these speeds are those of LM Studio or Ollama.
Questions about running an LLM locally
Which LLM for a Radeon RX 7900 XTX?
With 24 GB, Qwen3 32B or Gemma 4 31B fit at 4-bit, and mixtures of experts such as gpt-oss-20b write very fast.
What fits in 16 GB of VRAM?
Models of up to 14 billion parameters at 4-bit, or gpt-oss-20b, with room for a long conversation.
Does Twoody run on Linux with an AMD card?
Not yet: the Linux version is in preparation. Leave your email on the platforms page to be told when it is ready.
The best local LLM, by computer
Try local AI on your Mac, for free.
Version 0.11.5 · macOS 13 or later · no account, no subscription. Download it, install a model in one click, and ask your first question — even offline.