Software 44999 Published by

Strata 0.1.41 brings two headline defaults: per-GPU request concurrency that lifts throughput up to 87% on a four-card R9700, and CPU-assisted prefill that trims short-prompt latency on single NVIDIA cards. The multi-GPU fix cuts 4K-to-32K latency roughly in half and erases the connection resets that plagued bursts of 30 to 40 clients. Other gains include a Windows speedup for oversize models, a faster tokenizer cache, and a notable Intel Arc A750 decode bump. It ships on the MIT license with most new flags held off by default until they prove out.



Strata 0.1.41 Adds Multi-GPU Concurrency and CPU-Assisted Prefill

The open-source C++ inference engine folds in a Pascal hotfix plus faster multi-GPU throughput and optional CPU-assisted prefill on short NVIDIA prompts.

Another Strata has been released today. The new update is now at version 0.1.41, and the biggest change is what happens when you spread a model across more than one GPU. Developer Niko1221 published the release on 8 October 2026. It also inherits yesterday's Pascal hotfix and layers a long list of speedups and fixes on top.

Screenshot_from_2026_10_03_17_08_02

Strata is an MIT-licensed C++ engine that runs the MoE model Qwen3.8-Flash-Next on consumer hardware. You get a one-click installer for Windows and Linux, an OpenAI- and Anthropic-compatible API on localhost, and optional image input. It targets NVIDIA, AMD, and Intel Arc cards, including multi-GPU layer-split setups. As of the release it had around 18.3k GitHub stars and 1.6k forks.

Two context notes first. The previous release, 0.1.40.4, was a same-day hotfix that only restored decode speed for Pascal GPUs, meaning the GTX 10-series and Tesla P40. It changed no answers, which is why 0.1.41 simply absorbs that fix and adds everything else. The author puts it in one line: "Faster on more than one GPU, faster short prompts on NVIDIA, about twice as fast prompts on Windows with little RAM, and a long list of fixes."

Multi-GPU concurrency is the real story

Here's what actually shifts. With a layer split across cards, an older build would push one request group through every GPU stage while the other cards waited around. Now each GPU stage takes on its own group of requests at the same time, and the switch --batch-groups auto is now the default.

The numbers Niko1221 measured are the good part. On a four-card Radeon AI PRO R9700 box, total throughput at 8 clients jumped from 88 tok/s to 165 tok/s. That's a +87% gain. Median finish time (p50) dropped from 17.4s to 9.3s. Push it to 40 clients and you still pull 153 tok/s, up 75%, and those connection resets that used to plague runs of 30 to 40 clients are gone entirely.

That's after the server started listening with a backlog of 256. Keep in mind that's a genuine change, not just a throwaway knob. Answers stay identical too. An 8-request identity test passes on 2, 3, and 4 GPUs, and you can turn the whole thing off with --batch-groups 1. The engine still only serves 8 slots at a time, so anything past 8 clients queues. Worth remembering that the very first version of this feature slowed a long request sitting alongside short ones by 39 to 47%. That regression is fixed now and measured right in the release notes.

It's a rather compelling set of numbers for the internals involved, though the payoff only exists if you've already got multiple cards doing layer splits in the first place.

CPU-assisted prefill for short prompts

The second default is more niche. While a short prompt is being read, the idle CPU now grabs a slice of the experts the GPU would otherwise stream itself, but only for chunks under 1,024 tokens. Picture an agent's tool result or a short follow-up question.

On a single NVIDIA card, a 512-token prompt now reads 28% faster on an RTX 3060 and 33% faster on a Tesla P100. At 1,000 tokens the gains are 22% and 25%.

There's a catch, and it's a real one: answers can differ slightly. First-token KL versus the feature off averages 0.004. Want byte-exact 0.1.40.x answers? Set STRATA_PREFILL_CPU_SHARE=0. This only kicks in on a single NVIDIA GPU without --batch slots, so layer splits, batched runs, and AMD paths stay untouched.

On an RTX 5070 with Q2_0 the gain is a thin 0 to 2% unless you shrink the expert cache, since the setup's automatic cache already holds most of what a short prompt needs. Shrink it with --expert-cache 1500 and it climbs to 15 to 27%.

The single-GPU prefill win is real, though it only helps short prompts. However, at the same time, most modern agents are chock-full of short tool calls and follow-ups, so the practical upside is larger than the raw numbers suggest.

Beyond the two defaults, the changelog has meat. The standout is a Windows fix for models bigger than your RAM. Batched expert reads plus dropping the memory-mapped view of experts.bin cut prompt time 45 to 52% on an RTX 5070. It started as a port of adonizm's earlier work and got revived by malloc32.

The tokenizer cache now encodes a resent 89K-token agent prompt in 0.07s, down from 0.32s, with identical ids. An Intel Arc A750 decode climbs from 9.8 to 12.0 tok/s (+22%) because the sampler no longer needs FP64 emulation. Automatic card order now drops the faster card last, and a prompt already sitting in page cache stops getting hit with a 32-thread SSD profile.

The fixes are mostly plumbing, but a few land. A deadlock with parallel: 2 is resolved. The full conversation cache now evicts old entries instead of dropping the snapshot. A request that prints nothing for 90s gets caught by a new stall watchdog that ends and restarts it. Malformed API bodies now return a 400 with a traceback in the log, and --api-key finally takes comma-separated keys like llama.cpp. On the AMD side, there's a documented 925 to 2,503 tok/s prompt method for the RX 7900 XTX plus tuning tables for a couple of newer chip families.

Not every change is a win, though. Several new switches flip behavior but change answers or cost a little speed, so they sit off by default. STRATA_STAGE_PIN=1 in particular got pulled after the release check found corrupted IQ3_S answers, likely a buffer reused before its copy finished. That kind of catch is exactly why you run the release check.

The author shipped a long list of optional flags, none of which touch the default. --peer-device shares prompts without P2P through pinned host memory, but it's off because answers shift bits. STRATA_GDN_CHUNKED=1 chunks DeltaNet recurrence in 32-token pieces, only helping cards with 128 or more SMs. STRATA_ROUTE_RESIDENT prefers cached experts but changes answers, and STRATA_SM70_TABLE=1 adds V100 decode kernels. That last one is the tell. The team isn't shipping V100 kernels by default until someone actually benchmarks one on real hardware. Sensible call.

Update with ./update.sh. Models and configs are left untouched; setup just swaps in 0.1.41. To lock in exact 0.1.40.x answers on a single NVIDIA card, add "STRATA_PREFILL_CPU_SHARE": "0" to the env block of your strata-<model>.json. Head here to the GitHub repo to pull the release and read the full changelog yourself.