Strata v0.1.40 adds Strix Halo support two days after v0.1.39
A two-day release cycle brings the open-source LLM engine to AMD's Strix Halo APUs, a much faster Linux cold start, and a batch of stability and security fixes.
Strata v0.1.40 is out, and it arrived just two days after v0.1.39. That cadence says more about the project's energy than the size of the change, but 0.1.40 is a dense release. Niko1221's MIT-licensed C++ engine now runs natively on AMD's Strix Halo / Ryzen AI Max APUs, starts up about 13 times faster on Linux, and closes a pile of crash, NaN, and security holes.
For the unfamiliar, Strata is a free inference engine with a simple pitch. It runs models like Qwen3.8-Flash-Next on consumer hardware, ships a one-click installer for Windows and Linux, speaks an OpenAI/Anthropic-compatible API on localhost, and can take image input if you want vision.
Strix Halo goes local
The headline addition has no counterpart in 0.1.39. The HIP engine now recognizes AMD's Radeon 8060S / 8050S (that's the gfx1151 chip, PCI 1002:1586) and treats its shared memory as a single pool. Setup sizes it, recommends the UD-IQ4_XS quant, and on Linux auto-picks HIP.
On a Ryzen AI Max+ 395 with 128 GB running UD-IQ4_XS, Strata posted medians of 1,293 prompt tokens per second and 53.8 output tokens per second at 8K context. Those hold up through 64K and 128K, roughly 6% faster on prompts and 4% on output than the earlier Strix Halo engine branch.
Keep in mind that those 6% figures are against Strata's own earlier Strix Halo work, not v0.1.39. The older build couldn't run the chip at all, so the comparison is real but narrow.
The Windows HIP zip now ships gfx1151 code, but it's experimental and untested on Windows Strix Halo hardware. If you're on Linux, you're in better shape. There's a dedicated docs/STRIX_HALO.md for a manual, no-root build.
A faster Linux, a harder core
The change most people will actually feel is the Linux cold start. Weights, the head, embeddings, the vision encoder, and the RAM copy are now read ahead of use. On an RTX 5090 with 32 GB, running Q2_0 at 262K context, ready-time dropped to about 70 seconds from roughly 920. That's a 13x jump. Turn it off with STRATA_READ_AHEAD=0 if you'd rather not wait longer.
Linux also gains an unbuffered file tier (O_DIRECT), the same rule Windows already used. On a 16 GB laptop with 30 GB of RAM, that pushed decode from 4.4 to 7.9 tokens per second and trimmed startup from 36 to about 20 seconds.
Multi-GPU is where 0.1.40 really pulls ahead of 0.1.39. A new resident-RAM mode on a layer split, credited to Francesco Albano, ran about 70 tokens per second versus the 32 to 64 you'd expect on a 32 GB machine. It also fixes a bug introduced in 0.1.39, where adaptive swaps copied data back from the wrong card's cache. Setup keeps resident-on-split only if your engine supports it, so it now needs 0.1.40.
On four GPUs the layer split scores every four-way placement instead of guessing, which changes automatic placement on 4-GPU rigs. Two- and three-GPU setups keep 0.1.39's behavior. Stage weights can be trimmed on later stages too, and on one 4-GPU rig that reportedly took prompt generation from 470 to 1,930 tokens per second. Author's number. Take it as a best case.
A big chunk of this release is just hardening. NaN handling in fused SwiGLU q8_1 quantizers now keeps their scale finite, a fully resident split stage no longer waits on doorbells that never ring, and a dead engine refuses the next request instead of silently misbehaving. Issue #879 couldn't be reproduced, so it stays open. By the author's own pre-release checks, though, default output stays byte-identical across all four quants, matching 0.1.39 exactly.
Security moved from descriptive to active. 0.1.39 basically pointed you at SECURITY.md. 0.1.40 refuses network image paths, caps URLs at 32 MiB, and stops another origin from naming a local file as a picture, adding the same check to /v1/responses too.
The author also deliberately rejected two PRs, one that would let any website drive a Manager page and another that would turn an API key into a shell on a LAN server. Saying no to a feature is its own kind of security work.
It's also worth observing how much of this release is honest about what it didn't finish. The author runs on a single RTX 5070 with no second card, and lists exactly what still needs testers: 2-GPU soak tests, the Turing/RTX 20 path at long context, the Intel Arc SYCL port, older Pascal and Volta cards, and that untested Windows gfx1151 code. If you own any of those, the release notes are an open invitation.
Strata is MIT-licensed and self-funded through Buy Me A Coffee. Head here to the Niko1221/Strata repo on GitHub for the full notes and downloads.
