Strata: Run a 125-billion-parameter model on a gaming PC
Strata v0.1.38 spreads a frontier-scale language model across your GPU, CPU, RAM, and SSD so it runs on an ordinary gaming machine. The open-source inference engine from developer "Niko1221" handles the Qwen team's Qwen3.8-Flash-Next model, a 125-billion-parameter mixture of experts, on a 12 GB graphics card with 32 to 64 GB of RAM.
Run a frontier model on a rig that barely clears the bar for modern gaming. That is the pitch, and as of the v0.1.38 release (October 3, 2026) the GitHub repo holds about 7,900 stars and 700 forks. There are 600-plus commits and 120-plus merged pull requests, which is an unusually fast pace for a project of this size. The whole thing is fully local: nothing leaves your machine, nothing reaches a cloud server, and there's zero per-token cost once the hardware is bought.
What Strata actually does
Rather than fight a hardware limit, the engine leans into every piece of silicon and memory in the rig. The Qwen3.8-Flash-Next model is quantized using ISTA-DASLab's GSQ-RCO technique. It is built as a mixture of experts: 24,576 of them, with only about ten activated per token. Strata treats it like a working kitchen. The GPU keeps the most-used experts cached and within reach, the RAM holds all of them at once, the CPU chews through the ones not sitting on the card, and the SSD stores a large n-gram lookup table used during prompt processing.
There is also a "guess and check" trick, which is what the maintainers call the MTP layer. A small draft layer predicts the next few tokens while the main model verifies them all at once, yielding roughly 1.6 to 1.8 times faster responses without changing what you get. Strata speaks an OpenAI-compatible API at http://127.0.0.1:8080/v1 and an Anthropic endpoint, so it plugs into existing chat apps and coding tools like Cursor and Claude Code. Optional image input is available too.
What is new in v0.1.38
The headline item is a security fix, and honestly it is the one you should care most about. Local servers on your own machine are not immune to tricks like DNS rebinding and cross-site requests. The basic idea is that a sketchy website can point its own domain at your loopback address, then fire requests at a running server as if it were trusted. Without a fix, that opens the door to issuing commands to Strata's local server from a malicious page.
The answer is blunt: with no API key set, the server now only responds to hosts it trusts, meaning localhost, 127.0.0.1, [::1], and LAN addresses. Everything else gets a 403. Clients that send no Origin header, such as curl and the OpenAI and Anthropic SDKs, are unaffected, so normal programmatic use keeps working as-is. Once you set an API key, that key decides access and the host check is effectively switched off.
The rest of the update is mostly speed and compatibility. Prompt processing picked up three optimizations, all verified to produce bit-identical output. On an RTX 5070 12 GB the gains run from a sliver within noise on 4K prompts to 6.4% at 128K for IQ2_XS. Strata's people are upfront about a trade-off: since PR #463, adaptive-tier expert copies no longer depend on their arrival timing, which can shift draft acceptance on long prompts. So a figure that hit 75.1 tokens/s on one Q2_0 prompt dipped to 70.9.
The 4-bit KV cache, set with --kv q4_0, now reads on tensor cores for NVIDIA RTX 30-series and newer, trimming a 32K prompt on Q2_0 from 18.8 seconds to 14.3. Strata stresses that precision barely dents. Against an FP16 KV reference over 2,000 teacher-forced positions, top-1 match rose to 95.1% at 32K from 93.6% with the old kernel. The default --kv int8 is left alone.
Owners of smaller cards get a win as well. A 6 GB GPU could previously refuse to start when the default VRAM reserve consumed too much room. The engine now lowers that reserve, down to 300 MiB, until the cache fits, and warns if the card ends up nearly packed. Still short, and startup stops with a clear message saying exactly how many megabytes are missing and what you could free. Not bad for anyone living on the edge of their hardware.
AMD users see a bump too. v0.1.38 adds hipBLASLt tables for gfx1201 and gfx1200, delivering roughly 1.5 to 2 times faster prompts, plus a fix for a prompt hang on gfx1201.
Performance Numbers
The maintainers published a verification section more thorough than most open-source projects put out, and it shows. They checked identical answers to v0.1.37 across all four quantizations, ran speed A/B tests with alternating pairs, added kernel parity tests, and fired a real 57K-token prompt at the Coder model. Tests ran in bunches: 287 in tools/setup and 197 in the server, including the new security checks.
Keep in mind that Strata targets consumer hardware on purpose. A 12 GB NVIDIA RTX 20/30/40/50 card (or a compatible AMD card) is the sweet spot, though 8 GB runs, slowly. You need at least 32 GB of RAM, with 64 GB recommended if you want every model size. Keep 70 to 80 GB free on an SSD, NVMe strongly preferred, and be on Windows 10/11 or Linux. NVIDIA drivers must be 580 or newer.
The installer handles the rest: Python, CUDA libraries, the engine itself, and a roughly 70 GB model download that you can pause and resume. Which model you run depends on your RAM. On 32 GB you get the Coder model, which is best at code but weaker elsewhere. 48 GB opens up the IQ2_XS or Q2_0 packs, 64 GB fits every size, and 96 GB or more leaves room for the largest versions and Unsloth's experimental 4-bit build.
How AMD cards actually perform
That covers the notes. The AMD-specific gains (hipBLASLt tables for gfx1201 and gfx1200, plus a hang fix on gfx1201) are where I would push, and I got to check them myself. Running Strata on Debian 13 with Liquorix Kernel, ROCm, and a Radeon 7900 XTX, it was honestly one of the smoother AMD runs I've had this year. Output stayed steady, hitting up to about 120 tokens per second on the good prompts and settling into a more realistic 75 tokens per second across a mix of queries.
It is not flashier than the top NVIDIA cards, but the numbers feel believable rather than aspirational, which matters given what comes next. The Radeon 7900 XTX is a gfx1201 part, so it sits squarely in the family v0.1.38 optimized for. You will pay a little for the extra VRAM management overhead, but the point of Strata is to work on the hardware you already own, and on that front the AMD path now works without much fuss.
Strata is a small but real step toward the idea that you might not need the cloud to run frontier-class models.
Head here to the GitHub release page.
