Strata v0.1.43 makes running a 125B model on a single GPU noticeably faster
Another update has been released to Strata today, and its entire point is speed: faster decode and prompt processing on NVIDIA, AMD, and Intel GPUs. The project claims up to 38% faster decoding on a 16 GB Radeon over the previous version.
The catch, as there usually is one, is that the answers are byte-identical to 0.1.42. This isn't a smarter model. It's a faster one getting out of your way.
What the numbers actually say
The headline table is split by vendor, which is fair, since a given optimization rarely travels between architectures. Decode gains top out at +21% to +38% on a 16 GB Radeon AI PRO R9700 running the IQ3_XXS quantization. The same card with 32 GB of cache and the IQ3_S preset earns +6% to +14% on decode, with prompt processing climbing as high as +27% on 64K-context prompts.
NVIDIA gets quieter treatment. On Linux, RTX 30-series cards see +5% to +11% decode gains, and the aging Tesla P100 pulls a similar figure. Both are Linux-only. The RTX 5070 — the very card the project uses as its reference build — shows no gain by default on Windows, with NVIDIA changes sitting behind an opt-in flag.
Intel lands a surprising spot in the lineup. On the Arc Pro B70, decode jumps 14% to 21% while prompts nearly halve: a 4K prompt drops from 8.5 seconds to 3.3, and a 16K prompt falls from 27.4 to 13.4. That prompt speed is the part worth filing under "actually useful."
What Strata is, before the speed stuff
Strata targets exactly one model: Qwen3.8-Flash-Next, a 125-billion-parameter MoE from Alibaba's Qwen team. The trick lives in the architecture. Each token activates only about 6 billion parameters, waking roughly 10 of the model's 24,576 experts rather than burning the whole stack on every word. That sparse footprint is what makes local execution plausible at all.
The engine then spreads the model across three tiers simultaneously. VRAM holds the hot stuff plus an expert cache — each extra gigabyte stores roughly 700 more experts. Page-locked system RAM holds all 24,576 experts, and the CPU computes an expert in place whenever the GPU doesn't have it cached. An SSD holds a ~28.8 GB n-gram lookup table read one step ahead. The CPU doesn't block on the GPU driver; it spins on a doorbell.
On top of that sits speculative decoding via the model's own multi-token-prediction head. A fast path guesses a few tokens, the full 48-layer pass verifies them, and the ones it accepts get written out. The project credits this with a 1.6 to 1.8x speedup. The result is byte-identical to greedy decoding.
The price of all this? Not your GPU. It's your RAM. Strata's "12 GB is enough" headline is really "64 GB of RAM is enough." The card is only a cache for roughly 6 to 16 percent of the experts. If you're keeping score, the memory crunch of 2026 has made 64 GB of DDR5 rather pricey, so budget for that before you budget for the card.
Can you actually run it?
The floor is a 12 GB+ GPU, a current Linux or Windows 10/11 install, and a substantial chunk of system RAM. The installer picks your quant from your memory: Coder needs 32 GB, IQ2_XS and Q2_0 want 48 GB, and IQ3_XXS or IQ3_S need 64 GB to run every size. The recommended default is IQ2_XS. You'll also want about 80 GB of free disk, and an SSD is strongly recommended since the model download alone is roughly 70 GB.
The license is worth reading before you build anything commercial. The engine is MIT. The Qwen3.8-Flash-Next base weights carry Qwen's own qwen-community-1.0 license, and quantizations carry their own terms. Because a quant is a derivative of the base weights, the community license governs downstream use.
If that all sounds reasonable for your setup, here's where to go. The release notes and the full technical paper are on GitHub, and v0.1.43 adds a ./bench.sh script that runs a fixed suite against your own install in roughly 3 to 10 minutes and writes a scrubbed, shareable report.
Head here to the release page to grab it.
