Open-source Strata engine makes running a 125-billion-parameter model on your gaming PC notably faster
Strata v0.1.42 ships with ~13–16% faster decode on NVIDIA and a batch of fixes for multi-GPU splits, Windows AMD, Intel Arc and tight-RAM systems.
Strata v0.1.42 went out today, and the entire release hangs on a single number. Niko1221, the lead developer behind the MIT-licensed C++ inference engine, says it runs the 125-billion-parameter model Qwen3.8-Flash-Next notably faster on consumer hardware. On one NVIDIA GPU or a layer split, that lands at roughly 13 to 16 percent. Cards whose CPU outruns their PCIe link can push toward 30 percent.
The whole premise is what makes this interesting. A 125B model normally lives on data-center servers with hundreds of gigabytes of VRAM. Your best gaming GPU tops out near 24 GB. So Strata spreads the model across the entire machine instead of cramming it onto one chip. Qwen3.8-Flash-Next is a mixture-of-experts model, meaning it's really a team of about 24,576 tiny "experts," and any given word only needs around 10 of them.
Hot experts stay in your GPU's VRAM. All of them live in system RAM at the same time. The CPU picks up the rest while the GPU is busy, and a large lookup table sits on your SSD. A tiny helper model guesses the next handful of words before the big model validates them. Same output, just faster.
As of this release the repo sits at about 20.5k GitHub stars and 1.9k forks. That's a big enough community to churn through 295 open issues and 306 pull requests in parallel.
The headline speed comes from two defaults that switched on in this version.
Two defaults that changed the speed picture
The first is tail skip, controlled by an env var called STRATA_ROUTE_TAIL_SKIP. It's now set to 7 on CUDA, so it's on unless you say otherwise.
The idea is that during a verification window, an expert that isn't in VRAM and that every token ranks seventh or lower usually contributes almost nothing. Strata now skips copying and computing that expert. The rest keep the router's weights intact, so you aren't losing quality across the board.
The benchmark tables in the release notes are dense, so let me extract the part that matters. On an RTX 3060 with 12 GB of RAM running the IQ3_XXS quantization, median decode jumped 15.6 percent over five interleaved rounds. An RTX A4000 saw 12.9 percent. A Tesla P100 landed at 14.6 percent.
It only kicks in when experts are actually missing from VRAM. That's why it does nothing on a fully resident setup or a layer split, and it's off on AMD and Intel. Keep in mind that.
The second default is subtler. Strata used to guess how to split missed experts between the GPU reading them over PCIe and the CPU using a one-time startup probe. Now, with no manual flag, it times the CPU pool over the first decode windows, about 10 seconds, and balances the two. The step is 0.05, and it adjusts at most three times.
This is where the 30 percent lives. On the same RTX 3060 the share moved from 0.55 down to 0.20, and decode climbed from 47.7 to 62.1 tokens per second. On a P100 the gain was just 1.1 percent. The A4000 saw no change at all. Not every card benefits, for what it's worth.
On the AMD side, five prompt-path switches are now on by default for gfx1200 and gfx1201 cards, which includes the Radeon AI PRO R9700. Residual and logits come out byte-identical to the old path at 4K and 20K tokens. The reporter on an R9700 with IQ3_S saw prompt gains of up to 1.8 percent at 32K, though a different config reportedly ran 8 percent and 4 percent faster.
And here's the bit that matters if you care about reproducibility: the defaults do shift answers slightly.
Strata measured this properly. Against the same engine with tail skip off, an RTX 3060 with IQ3_XXS agreed on 97.7 percent of tokens in greedy mode. The A4000 with Q2_0 hit 98.2 percent. For scale, running IQ3_XXS versus Q2_0 on the same text differs by about 0.10, so a fraction of a percent here won't derail most people. However, at the same time, if you're pinning exact outputs for a benchmark, keep in mind that setting STRATA_ROUTE_TAIL_SKIP=0 along with the old prefill and PCIe flags gives you exact 0.1.41 answers.
Fixes, opt-ins and what's still broken
Beyond speed, this release quietly folds in five fixes that were originally slated for a 0.1.41.1 hotfix. Those never shipped separately, and they're now baked in.
One of them fixed something that had already bitten a reporter. An automatic layer split used to put the faster card last, and on unequal-VRAM pairs that stashed the head, draft and verify steps on the weakest card. The reporter's 12 GB card dropped from roughly 58 tokens per second to 44. Strata now skips the reorder when it would dump the last stage on a card with less total VRAM, and cards within 5 percent of each other keep your order.
There's also a fix for Windows AMD where a bare prefill mode stalled at 6,144 tokens and the process died with a heap-check error, costing about 14 percent prompt speed. The usual spread across Intel Arc, Pascal and Volta with little free VRAM is handled too, along with a check that a 4,094-token prompt doesn't bail out with a misleading "ran out of context" message.
The most useful might be the simplest. Strata now handles CPUs without AVX2, which previously killed the build with a SIGILL. It also honors container memory limits instead of guessing at the host's total RAM.
If speed is really the point, the release also throws in a pile of opt-in speedups that don't touch the default behavior.
The most eye-catching is data-parallel replicas. Run N engines behind one server, each on its own card grouping or set of cores, and a conversation bounces back to the same replica. On four R9700s with eight clients, throughput went from 130 to 160 tokens per second. Forty clients ran at 1.39x. The catch is that you load two copies of the model into RAM, and core pinning is Linux only.
Fused int8 prompt experts on gfx12 stand out too, pushing R9700 prompt throughput from 761 to 876 tokens per second. It stays opt-in because the first-token accuracy drift is too high for the tighter quantizations.
The list keeps going: pipeline windows on three or more layer-split stages, --batch-mtp on a split, a couple of routing switches that change answers, Intel A-series prompt kernels, session-reclaim flags and a Prometheus metrics endpoint behind the API key.
It's arguably a lot of knobs for a one-click installer, but Strata was never aiming to be minimal. The design exists to squeeze the most out of hardware that's fundamentally too small for the model.
Some things still didn't get fixed. An Arc A750 can't run --batch without tripping an out-of-host-memory error. A Windows AMD setup with KV streaming still throws an unspecified launch failure, though Strata can't even test that combination in-house, so those read as guards that compile rather than proven fixes.
Open issues remain, including a Windows crash while parking a very large conversation. And yes, that's the same failure class the card-order fix was working around.
Grab the release with ./update.sh on Linux. Your models and configs stay untouched since setup swaps the engine in place.
Head here to the Strata releases page for the direct binary if the scripts don't do it for you. To get the exact old answers back on NVIDIA, drop three env vars into the "env" block of your strata config and you're done.
