Strata Speeds Up Local Qwen and Adds OpenAI Codex Support
Version 0.1.39 improves local inference and developer integration, but its performance gains depend on memory, quantization, and the hardware path you choose.
Listen
AI narration
15:03
0:00 / 15:03
AI SummaryGenerated from this article
Strata v0.1.39 accelerates local Qwen inference with 2.5% to 7% faster decoding on RTX 5070 and adds OpenAI Responses API support for Codex CLI integration. Performance gains depend on quantization, hardware configuration, and memory availability. Optional parallel serving reduces request queues but trades throughput for responsiveness. The new API endpoint lets Codex clients connect to local Qwen3.8-Flash-Next without workflow changes, though the implementation omits hosted tools and reasoning summaries. Experimental paths extend support to older NVIDIA, AMD, and Intel GPUs.
Strata v0.1.39 makes running Qwen3.8-Flash-Next locally more practical with faster inference on documented consumer-GPU configurations and an OpenAI Responses API endpoint that lets Codex CLI connect to the local model.
The project reports roughly 2.5% to 7% faster decoding in its RTX 5070 tests against v0.1.38, plus a 26% gain processing a 243,000-token prompt on an RTX 3060. Optional parallel serving and experimental older-GPU paths broaden what builders can do with it.
These are project-reported measurements, not independently established performance guarantees. More concurrent requests can make a memory-constrained PC slower, and fitting a heavily quantized model locally doesn't establish equivalent answer quality.
Faster Decoding, With Specific Benchmark Boundaries
The v0.1.39 release notes, dated October 4, 2026, describe changes that reduce GPU launches and CPU-GPU round trips during decoding, the stage that generates an answer after the prompt has been processed.
The maintainer measured the following results on one RTX 5070, comparing v0.1.39 with v0.1.38 over 10 interleaved pairs and reporting medians:
Quantization
Workload
v0.1.38
v0.1.39
Q2_0
Story
68.4 tokens/s
72.7 tokens/s
Q2_0
Code
75.8 tokens/s
80.0 tokens/s
IQ3_XXS
Story
46.1 tokens/s
48.8 tokens/s
IQ3_XXS
Code
50.7 tokens/s
51.9 tokens/s
These are modest but useful gains without changing the selected quantization. The maintainer also reports byte-identical output in fixed-cache comparisons across four quantizations, within the tested samples.
Long-prompt improvements are less uniform. Strata changes how it sizes its streamed-expert buffer and selects prompt-processing chunks. On an RTX 5070, a 32K-token IQ3_XXS prompt improved by 18.5% with a 1,500-slot expert cache, but showed no improvement with the installer's default configuration.
An attention-kernel change delivered a reported 26% gain for a 243K-token prompt on an RTX 3060. That is a specific processing improvement; it doesn't establish that every consumer PC can handle such contexts comfortably.
Work with Zeniteq
Let’s work together
We’re open to thoughtful collaborations with teams building in AI. Explore the ways we can work together.
There are regressions, too. The Coder configuration processed a 30K-token prompt about 8% slower when configured for a 32K context, although it improved at a 64K context. STRATA_RING_BYTES=0 restores the previous streamed-expert buffer behavior.
Long-prompt output can differ from v0.1.38 because computation follows different cached and streamed groups. The project reports similar quality in its limited teacher-forced comparison against an FP16 prompt-processing path. That check is useful, but it isn't an end-to-end coding or reasoning evaluation.
OpenAI Responses API Support Connects Codex to Local Qwen
The new POST /v1/responses endpoint addresses a concrete integration problem: clients that expect OpenAI's Responses API need more than a generic chat endpoint.
According to the release notes, Strata's implementation supports function tools, tool results, reasoning-effort settings, JSON schemas, and streaming events. The maintainer tested it with Codex CLI 0.160.0, including a tool loop, and reports that later turns reused approximately 96% of the prompt from cache in that test.
Codex supplies the client workflow; Strata serves Qwen3.8-Flash-Next as the underlying model. The connection provides API compatibility, not an OpenAI model running locally. Tool selection, coding accuracy, and instruction following still depend on Qwen and its configuration.
The implementation is stateless and does not support previous_response_id, hosted tools, or reasoning summaries. Clients expecting those features cannot assume complete Responses API compatibility.
Developers already using Codex can try a local backend without changing coding clients. The Strata repository links its configuration instructions and documents the existing OpenAI-compatible localhost interface.
Parallel Serving Reduces Queues, Not Necessarily Runtime
Strata still serves one request at a time by default. Version 0.1.39 adds optional concurrent decoding through "parallel": N in the model configuration, or setup --parallel N.
Up to the configured number of conversations can run together. Additional requests wait for a slot, and a long prompt yields to a short one at a processing-chunk boundary.
The project’s RTX 5070 Q2_0 test illustrates the tradeoff. With four requests running concurrently, the last request began after 1.8 seconds instead of 11.2 seconds. Aggregate decoding, however, was approximately 11% slower: 63.1 versus 70.7 tokens per second.
Even a request running alone suffered when extra slots were configured. The maintainer reports an 11% loss with two slots and 22% with four because each session reserved 0.56 GiB that otherwise could serve the expert cache.
Shorter waits can make an interactive service feel more responsive despite slower generation. A single-user coding agent may be better served by the default. The project recommends concurrency primarily where the experts mostly fit in VRAM, not as a universal throughput upgrade.
Consumer Hardware Still Needs Substantial Memory
Strata is free, MIT-licensed software for Windows and Linux. Its documented standard requirements include a supported NVIDIA or AMD GPU with at least 12 GB of VRAM, 32 GB of system RAM, and approximately 80 GB of free disk space. Experimental paths have separate requirements.
The engine combines system-memory storage, GPU expert caching, and streamed computation to run a model that cannot fit entirely into a typical gaming card. Quantization, which stores model weights at reduced precision, is central to making that possible.
The repository recommends the Coder variant for 32 GB systems, smaller quantizations for 48 GB, and offers more choices at 64 GB. Larger versions can require considerably more memory. Those choices also affect quality: the project describes Coder as a coding-oriented version with half the experts removed and warns that it is weaker outside coding, including Chinese and other CJK text.
Higher-quality configurations can become storage-bound. The README reports only 7 to 8.5 tokens per second for its experimental UD-Q4_K_XL configuration on a 64 GB PC because most weights are read from SSD during generation.
Version 0.1.39 fixes a RAM-budget regression that caused unnecessary drive reads. It doesn't eliminate the underlying cost of insufficient memory.
Older GPUs Gain Experimental Routes
The release expands hardware access while explicitly distinguishing community-contributed paths from hardware the maintainer tested locally.
For NVIDIA Pascal and Volta cards, including the P40, P100, GTX 10 series, V100, and Titan V, setup can select a separate CUDA 12.9 engine. The release notes explain that CUDA 13 cannot compile for those architectures.
That path can also accommodate certain older NVIDIA drivers. It is not the recommended default for current cards, and the maintainer explicitly lists Pascal, Volta, and old-driver configurations as untested locally.
Other experimental routes include:
Older AMD hardware such as Radeon VII and Instinct MI50/MI60, with manual builds for some architectures.
RX 6700 XT support through setup.
An Intel Arc SYCL backend built from source on Linux, without a ready-made Intel engine.
Older CPUs without AVX2, with slower fallback builds and restrictions on supported quantizations.
The maintainer says these paths were contributed and measured by community members. Compilation and unit tests alone do not establish reliable performance across every listed device.
For owners of spare hardware, the expanded routes are worth investigating. They shouldn't be read as a blanket promise that an old workstation will match a supported modern gaming PC.
Quantization, Vision, and Setup Still Need Scrutiny
The Hacker News discussion showed enthusiasm alongside practical objections. Those comments are individual reports, not representative testing of Strata's user base.
Commenter a11r questioned the quality implications of going below four-bit quantization, while ricardobeat described mixed results with IQ3_XXS. Other participants reported positive coding experiences. Given that disagreement, a chosen quantization needs testing on actual workloads; generation speed alone doesn't establish usefulness.
Vision deserves a separate evaluation. Commenter Jackson__ reported worse object-coordinate accuracy through Strata than through llama.cpp using the same model and vision-adapter weights. This claimed comparison from one commenter is not a verified general finding, though it identifies a specific capability to check before relying on image input.
The repository also documents platform limits: AMD image processing uses the CPU on Linux, while AMD image input is not yet supported on Windows.
Some commenters raised concerns about copy-pasted setup commands. Strata's README offers both direct installation and instructions that delegate setup to an AI coding assistant. In either case, users still need to review what an installer or agent will execute.
The release documents localhost binding at 127.0.0.1 by default, API-key handling, and Host and Origin checks. Its network-access instructions tell users to set a secret API key when binding beyond localhost. These are documented controls, not evidence of an independent security audit.
Final Thoughts
The most defensible reason to adopt v0.1.39 is that it removes practical friction around an existing local model: a compatible Codex client path, shorter queues when memory permits, and measured improvements without requiring a new GPU.
Judge the update by completed work. Compare the same tasks, quantization, and context settings before and after updating, including tool-call correctness and image accuracy if those matter to the workflow. A faster backend is useful only if the chosen configuration retains enough quality to avoid spending the saved time correcting its output.
Frequently Asked Questions
3 questions
1
Does Strata v0.1.39 Work With OpenAI Codex?
Strata v0.1.39 adds a Responses API endpoint tested with Codex CLI 0.160.0, including a tool loop. Codex can use Strata as a local backend, but the underlying model is Qwen3.8-Flash-Next, not an OpenAI model. The implementation does not support hosted tools, reasoning summaries, or previous_response_id.
2
Will Parallel Serving Make Strata Faster?
Parallel serving can reduce waiting time without improving generation throughput. In the project's RTX 5070 test, four concurrent requests started sooner but decoded approximately 11% slower overall. Additional sessions also reduced the expert-cache memory available. The default single-request mode may be preferable on memory-constrained PCs.
3
Can Strata Run on Older GPUs?
Strata v0.1.39 provides experimental paths for older NVIDIA and AMD GPUs, plus a source-built Intel Arc backend on Linux. Support varies by architecture and may require a separate engine or manual compilation. The maintainer explicitly says several paths were not tested locally on the target hardware.