Strata v0.1.39 makes local AI faster without changing a single byte of output
Strata, the MIT-licensed open-source LLM inference engine from developer Niko1221, shipped version 0.1.39 today. The release doesn't arrive with a flashy feature. Instead it keeps the exact same answers as 0.1.38 while decoding faster, handling longer prompts, juggling multiple conversations, reaching older hardware, and speaking OpenAI's Responses API so Codex CLI can drive a local instance.
The version number sounds trivial. 0.1.39 up from 0.1.38 is barely a tick, and that's almost the point. This reads like a mature project polishing a core promise rather than chasing a launch moment. The repository currently sits at roughly 10,000 stars and 888 forks, with 68 open issues and 156 pull requests in flight.
Strata runs Qwen3.8-Flash-Next, a 125-billion-parameter model that normally needs server hardware, on an ordinary gaming PC with 12 to 24 GB of VRAM. Nothing leaves your machine. Chat, code generation, image understanding, and agent integration all happen locally.
What Strata actually runs
Qwen3.8-Flash-Next is a strange model by design. It carries 125 billion total parameters but only activates about 6 billion per token, using a Mixture-of-Experts layout where each word triggers just a handful of experts. A distinctive 51-billion-parameter n-gram layer lets it pay for scale in memory instead of compute. Add a 4-billion-parameter draft layer that guesses upcoming tokens, and generation runs on a guess-then-check loop that's roughly 1.6 to 1.8 times faster than a straight pass.
The numbers get impressive from there. 48 layers, a hidden dimension of 2,560, and native 262,144 tokens of context, extendable to roughly a million with YaRN. On Qwen's own benchmarks it matches Qwen3-27B on agentic coding and approaches Claude Opus 4.6 on reasoning, landing 91.7 on GPQA Diamond.
The engineering trick is squeezing that onto hardware that "usually runs on servers with hundreds of gigabytes of graphics memory." Strata uses a kitchen analogy: common ingredients stay on the counter, the rest waits in the pantry. The GPU holds attention mixers, shared experts, the output head, and an adaptive cache that packs the most-used experts into leftover VRAM. RAM stores all 24,576 experts. The SSD holds a 28.8 GB n-gram table read a few rows per token.
Faster, and provably unchanged
The headline change is a faster decode, driven by PR #646 from Stuart Chapin. The verify pass now launches fewer kernels and makes fewer host round trips. On a single RTX 5070, decoding jumped about 6% on Q2_0 weights, pushing story generation from 68.4 to 72.7 tokens per second and code from 75.8 to 80.0. IQ3_XXS gained 6% and 2.5% across two test sets.
Here's the part that matters. Output is byte-identical to 0.1.38 across all four quants they tested. Not close. Identical.
The gain only activates when every expert of a layer fits in VRAM, which a 12 GB card never reaches, so the biggest gains weren't measured here. That's an honest caveat.
Longer prompts now ship by default under PR #583. The streamed expert ring is sized in bytes, and an automatic prefill picks the largest full chunk. On an RTX 5070 with 32K prompts, IQ3_XXS gained about 18.5% with a 1,500-slot cache. But Strata admits the bits genuinely change versus 0.1.38 for long prompts, since experts pass through a different cached mix. Quality stays in the same band anyway: KL divergence of 0.020 versus 0.019, top-1 accuracy 95.3% versus 95.0%.
It doesn't help everyone. The Coder model reads a 30K prompt about 8% slower at 32K context, though it flips to ~6% faster at 64K. You can restore old behavior with STRATA_RING_BYTES=0.
New capability: multiple GPUs and parallel requests
Multi-GPU support is opt-in. Add --remote-expert-opt to a two-or-more GPU config and Strata skips host work for tokens whose experts live entirely on helper cards. A dual RTX 4090 setup measured 63% faster on mixed text and 132% faster on code. It's opt-in because the rounding differs, so one or two GPUs keep running exactly as before.
Strata can also answer several requests at once now. The old 0.1.38 handled just one. PR #465 adds a parallel option that decodes up to N conversations together, with extra requests waiting for a free slot. On a 12 GB card, four requests dropped the last one's start time from 11.2 seconds to 1.8 seconds. The catch is the batch decodes 11% slower overall, since each slot burns 0.56 GB of the expert cache.
So batching only makes sense when experts mostly fit in VRAM. A four-GPU layer split served eight requests at 360 tokens per second versus 120 for one.
Codex CLI, and older hardware that modern engines abandoned
The piece probably most relevant to developers. Strata now serves OpenAI's Responses API under PR #451, so Codex CLI works with it. Tested with Codex 0.160.0 including a tool loop, with later turns reusing about 96% of the prompt from cache. Not supported are previous_response_id, hosted tools, and reasoning summaries. Details for the Codex config live in docs/DETAILS.md.
This release also extends life to gear the main engine ignores. Older NVIDIA cards like the Pascal and Volta family, the P40, P100, GTX 10, V100, and Titan V get a second CUDA 12.9 build kept around, since CUDA 13 can't compile for them. It activates only when you select such a card. AMD gets hand-built paths for gfx906 and gfx1012, Intel Arc picks up a SYCL port, and CPUs without AVX2 now compile instead of bailing.
Keep in mind that Strata compiles and unit-tests these paths but has none of the hardware itself. So the phrase "untested here" shows up a fair amount for Intel Arc and AMD-on-card. Reported figures like V100 at 1,123 to 1,251 tokens per second are best treated as rough.
Is it worth your time?
It's free software, so the real cost is your hardware and setup time. For anyone with an NVIDIA RTX 20 series or newer and 12 GB or more of VRAM, this is a genuinely serious local option, and arguably one of the best free ones right now. The byte-identical correctness claims and the sheer volume of validation are what stand out. A release that tests kernel parity, runs a 59K-token prompt, and reports 376 setup tests plus 268 server tests is unusual discipline for the open-source AI space.
The experimental paths are genuinely experimental, and that's fair. Older cards, Intel Arc, and AMD-on-card all carry the caveat, which is at least honest about limits.
Head here to the GitHub release page for the full changelog and the two companion documents in the repo,
