Software 44984 Published by

Strata is a free, open-source C++ inference engine that runs the ~125-billion-parameter Qwen3.8-Flash-Next model on consumer gaming PCs without sending data to the cloud. Its new v0.1.40.2 release, out today, officially adds Intel Arc GPU support on Linux while lifting speed across NVIDIA and AMD hardware. The biggest gains show up in multi-GPU splits, machines with less RAM than the model, and long-document workflows, and default output stays byte-identical to 0.1.40. Kept on GitHub by Niko1221, the MIT-licensed project keeps shipping multiple releases per week, leaning on community contributions to test hardware it can't own itself.



Strata v0.1.40.2 adds official Intel Arc support and wider local inference gains

Run a 125-billion-parameter language model on a gaming PC without sending a byte of your data to any cloud. Strata does that. With v0.1.40.2, released today, it also got faster across NVIDIA and AMD hardware and, more importantly for Intel owners, finally started working on Arc GPUs.

Strata is a free, open-source inference engine written in C++. It runs Qwen3.8-Flash-Next, a ~125B-parameter mixture-of-experts model that normally needs data-center hardware, on an ordinary consumer desktop. It's maintained by a developer who goes by Niko1221, based in Ljubljana, Slovenia. The GitHub repo carries roughly 16,600 stars and 1,400 forks. It's MIT licensed, so you're free to dig in.

Screenshot_from_2026_10_03_17_08_02

The pitch is partly privacy. Nothing leaves your PC. A single click installs it on Windows or Linux, then it talks to your existing software through a standard OpenAI-compatible API, an Anthropic-style endpoint, and the OpenAI Responses API for tools like Codex CLI. There's a browser app at http://127.0.0.1:8080, with chat, a live hardware monitor, and settings.

The trick to fitting such a giant model on modest hardware is how it splits the work. Qwen3.8-Flash-Next has roughly 24,576 small "experts" and only needs about 10 per word, so Strata never keeps the whole thing on the graphics card. The GPU holds the hot experts plus attention and routing. The system RAM holds all the experts, and the CPU works through the ones not currently on the GPU at the same time, so neither side idles. A large n-gram table lives on the SSD, read only a few rows per token.

There's also a "guess and check" step. An internal draft layer guesses up to three next tokens, then a single pass over the 48 layers verifies them. That nets roughly 2.4 to 3.2 tokens per pass at no quality hit, which the project values at about 1.6 to 1.8x speed.

On supported hardware (NVIDIA RTX 20/30/40/50 series or selected AMD Radeon cards with 12 GB+ VRAM, and 32 to 64 GB of RAM), Strata can write at 50 to 90+ tokens per second. That's faster than you can read, and it swallows long prompts at over 1,000 tok/s.

v0.1.40.2 frames itself as "faster and fixed across the board." Three themes: official Intel Arc support, a pile of bug fixes pulled from open issues and PRs, and real speed gains in multi-GPU, low-RAM-on-Linux, and long-document cases.

One reassuring detail: default answers are byte-identical to 0.1.40, verified across the Q2_0, IQ3_XXS, Coder, and IQ3_S quantizations. In other words, this changes how fast Strata runs, not what it says. Keep in mind that last part, since several of the fastest features are opt-in precisely because they change the output.

Raw decode gains on a single card were modest. Strata measured +3.9% on an RTX 5070 with IQ3_XXS and +4.3% on an RTX 3060 for code. Small, but consistent.

The bigger numbers show up in specific setups. A multi-GPU layer split, fixed when the automatic choice went unbalanced, can gain up to 60% prompt speed. On four Radeon AI PRO R9700 cards, a 32K prompt now reads at 2,797 tok/s. Low-RAM Linux machines benefit most when the RAM is smaller than the model itself: on a Tesla P100 capped at 32 GB, story decode jumped from 12.6 to 15.2 tok/s and code from 9.3 to 13.3, with three to seven times less drive traffic.

Long-document workflows see the most dramatic jump. With an opt-in pin=N, the first answer can land in 0.35 to 0.55 seconds instead of 48 to 149, because a shared prefix is read once and kept. And on a GB10 with --mmap-experts, IQ3_S prompt throughput went from 91 to 1,222 tok/s.

Official Intel Arc support

This is the headline for Intel owners. 0.1.40 didn't build for Intel at all. Now Strata's engine runs on SYCL, sitting behind the same server, APIs, and web app used by NVIDIA and AMD. Installation runs python3 sycl/setup_intel.py, which builds the oneAPI stack inside a Docker image.

It was tested on an Arc Pro B70 (32 GB) and an Arc A750 (8 GB). Both passed a 10-prompt correctness check with multi-token prediction, and neither hung. The Arc Pro B70 hit roughly 980 to 1,000 tok/s on a 4K prompt with IQ3_S, decoding at 31 for story and 41 for code. The A750, with less memory, sat around 58 tok/s prompt and 16 tok/s decode. That's not going to blow anyone away, for what it's worth, but it at least works.

It was a real slog to get there. The port had to deal with the xe driver returning aliased GPU memory, hangs when the GPU got pageable host memory, stalled draft-token acceptance, and first-window hangs on older i915 A-series cards. Each got checks, reallocation loops, and queue-drain fixes. Arc support stays Linux-only for now.

Bug fixes and housekeeping

The rest is a broad sweep. Strata recently cleaned up its git history, so the update scripts now auto-migrate an old clone, preserving commits on a pre-cleanup-backup branch. There's a fix for a host out-of-memory crash at start on Linux when RAM was mostly file cache, and the server now bounds its hangs with timeouts of 900 seconds for the engine and 300 for the encoder.

Other patches cover garbage output from STRATA_QFUSE, hung batch slots on multi-GPU setups, and doubled numbers in long KV-count displays. Setup got hardened too: ROCm wheels are now chosen per GPU family, CUDA 13.2.2 is accepted, and downloaded engines are verified against a SHA-256 digest. Smaller touches include a --gpu LIST to pick visible GPUs, a Prometheus /metrics endpoint, EXIF orientation handling for uploaded images, and 14 new benchmark reports.

Release builds still use CUDA 13.0, matching 0.1.40. To update from 0.1.40.1 or newer, run ./update.sh as usual. From 0.1.40 or older, do one manual git step first, then the script. Models, configs, and the engine are left untouched.

For context, this is the third release in a very short window. v0.1.40 landed two days earlier on 6 October, adding Strix Halo support, new quantization formats, and a faster Linux cold start. v0.1.40.1 was a Python-server hotfix. v0.1.39 brought concurrent request handling and the Responses API for Codex CLI.

Multiple releases a week is a tell that this is a community-driven project moving fast. Niko1221 openly asks for testers on hardware it can't personally own, like multi-GPU rigs, and credits contributors whose forks feed back in. It stands as part of a broader push toward running models that once needed cloud infrastructure entirely on personal hardware.

Strata is MIT licensed and sustained by optional donations. Head here to the GitHub repository to grab v0.1.40.2 and read the full release notes.