Twoody

Local LLM calculator

The best local LLM for your AMD graphics card

From the RX 7600 to the RX 9070 XT: the strongest model that fits each card's memory, and how fast it writes.

Your computer Memory Recommended — runs comfortably Writes

Radeon RX 9070 XT (16 GB)

16 GB gpt-oss-20b 100+ tokens/s

Radeon RX 9070 (16 GB)

16 GB gpt-oss-20b 100+ tokens/s

Radeon RX 9060 XT (16 GB)

16 GB gpt-oss-20b 61–107 tokens/s

Radeon RX 9060 XT (8 GB)

8 GB Llama 3.1 8B 42–61 tokens/s

Radeon RX 7900 XTX (24 GB)

24 GB GLM-4.7-Flash 100+ tokens/s

Radeon RX 7900 XT (20 GB)

20 GB gpt-oss-20b 100+ tokens/s

Radeon RX 7900 GRE (16 GB)

16 GB gpt-oss-20b 100+ tokens/s

Radeon RX 7800 XT (16 GB)

16 GB gpt-oss-20b 100+ tokens/s

Radeon RX 7600 XT (16 GB)

16 GB gpt-oss-20b 46–80 tokens/s

Radeon RX 7600 (8 GB)

8 GB Llama 3.1 8B 36–52 tokens/s

Estimates, not measurements, computed from public llama.cpp benchmarks. Figures reviewed on September 26, 2026.

The card's memory decides which models fit: 16 GB runs models of 14 billion parameters or gpt-oss-20b; the 24 GB RX 7900 XTX, 32 billion.

llama.cpp runs Radeon cards through Vulkan: writing is fast; reading long documents is slower than on NVIDIA cards.

Twoody for Windows and Linux is in preparation; the version being tested runs models on the processor. Meanwhile, these speeds are those of LM Studio or Ollama.

Questions about running an LLM locally

Which LLM for a Radeon RX 7900 XTX?

With 24 GB, Qwen3 32B or Gemma 4 31B fit at 4-bit, and mixtures of experts such as gpt-oss-20b write very fast.

What fits in 16 GB of VRAM?

Models of up to 14 billion parameters at 4-bit, or gpt-oss-20b, with room for a long conversation.

Does Twoody run on Linux with an AMD card?

Not yet: the Linux version is in preparation. Leave your email on the platforms page to be told when it is ready.

Try local AI on your Mac, for free.

Version 0.11.5 · macOS 13 or later · no account, no subscription. Download it, install a model in one click, and ask your first question — even offline.