● GeneratorNest

Strata Qwen 3.8: Run a 125B Model on a 12 GB GPU

Updated 2026-10-09

Short answer: Strata is a new open-source (MIT-licensed) inference engine, published on GitHub in early October 2026, that runs Qwen3.8-Flash-Next – a 125-billion-parameter mixture-of-experts model – on a single consumer GPU with 12 GB of VRAM plus 32 GB of system RAM. Its own numbers report about 94 tokens per second of output on an NVIDIA RTX 5070 and about 60 on an AMD RX 9070 XT at Q2_0 quantization. These are the project's self-reported figures; no independent benchmarks had been published as of October 4, 2026.

That is why "strata qwen 3.8" suddenly became a search: a model that normally needs datacenter hardware, running on a gaming PC. Here is what is going on and what you need to try it.

How can a 125B model fit on a 12 GB card?

The trick is the model's architecture, not magic compression alone. Qwen3.8-Flash-Next is a mixture-of-experts (MoE) model. Strata's README explains it plainly: the model is "a team of 24,576 small specialists ('experts')" and "each word needs only 10 of them." Only a tiny part of the 125 billion parameters is active for any one token.

Strata uses that by splitting the model between two kinds of memory:

  • Hot experts on the GPU. The experts the model uses most often stay in the card's fast VRAM.
  • Everything else in system RAM. The rest is offloaded to normal memory and pulled in when needed. The README compares it to a kitchen: "the things you use all the time stay on the counter, and the rest waits in the pantry."

On top of that, Strata includes a speculative drafter – "a small helper guesses the next few words. The big model checks them all at once" – which the project says adds another 1.6–1.8x throughput.

Reported speeds

All figures below are Strata's own, as reported on October 4, 2026:

GPU Quantization Output speed Input (prompt) speed
NVIDIA RTX 5070 (12 GB) Q2_0 ~94 tokens/s ~2,650 tokens/s
NVIDIA RTX 5070 (12 GB) IQ3_S ~53 tokens/s –
AMD RX 9070 XT Q2_0 ~60 tokens/s ~1,160 tokens/s

The trade-off is visible in the table: the more aggressive Q2_0 quantization is fastest, while IQ3_S keeps more quality at roughly half the speed. Q2 quantization of any model loses some quality compared with full precision, so test it on your own tasks before relying on it.

What you need

  • A GPU with at least 12 GB of VRAM. The project's examples use an RTX 5070 and an RX 9070 XT; guides written since then also show it on an RTX 4090.
  • 32 GB of system RAM or more, because most experts live there.
  • Fast storage for the quantized weights (tens of gigabytes even when compressed).
  • Comfort with the command line. Strata builds on llama.cpp / ggml, and the compressed weights credit ISTA-DASLab, UkisAI and Unsloth.

How to try Strata with Qwen 3.8 (outline)

  1. Check the official repository (Strata on GitHub, MIT license) for the current install steps – early projects change quickly.
  2. Download the quantized Qwen3.8-Flash-Next weights the README links to (start with Q2_0 if you have exactly 12 GB of VRAM).
  3. Run the engine with the default GPU/RAM split, then adjust how many experts stay on the GPU if you have spare VRAM.
  4. Turn on the speculative drafter and compare tokens per second with and without it on your own prompts.
  5. Compare quality between Q2_0 and IQ3_S on a few real tasks before choosing.

Because nothing has been independently reproduced yet, treat your first runs as a benchmark of your own machine, not as a guarantee of the published numbers.

Is it worth it?

Worth trying if you want a large model running locally for privacy, offline use or tinkering, and you already have a 12 GB+ GPU and 32 GB of RAM.

Probably not yet if you need dependable production quality: heavily quantized weights, a brand-new engine and self-reported benchmarks are a lot of unknowns at once. For everyday writing tasks, a hosted model is still the simpler option – for example our free AI text tools.

FAQ

What is Strata? An MIT-licensed open-source inference engine, released in early October 2026, built to run large mixture-of-experts models such as Qwen3.8-Flash-Next on a single consumer GPU by keeping frequently used experts on the GPU and the rest in system RAM.

Can Qwen 3.8 really run on a 12 GB GPU? According to the project, yes – Qwen3.8-Flash-Next (125B) with 12 GB of VRAM and 32 GB of RAM, using Q2_0 or IQ3_S quantization. Independent benchmarks were not available at the time of writing.

How fast is Strata? Self-reported: about 94 output tokens per second on an RTX 5070 and about 60 on an RX 9070 XT at Q2_0, plus a claimed 1.6–1.8x boost from speculative decoding.

Is Strata made by the Qwen team? No. Strata is an independent open-source project; Qwen3.8-Flash-Next is the model it runs.

Does it work on Mac or CPU only? The published results are for NVIDIA and AMD GPUs. Check the repository for current platform support.

Try it free

100% free – no sign-up needed.

Open the tool