A GitHub project called Strata claims it can run Qwen3.8-Flash-Next, a 125-billion-parameter model, on a gaming PC with 12 GB of video memory. The repo publishes measured throughput on two ordinary machines: an RTX 5070 (12 GB) paired with a Ryzen 5 7600 and 64 GB of RAM hits 94 tokens per second at the Q2_0 quantization and 53 tokens/s at IQ3_S, while an RX 9070 XT (16 GB) with a Ryzen 9 3900X and 47 GB of RAM reaches 60 tokens/s at Q2_0. Prompt ingestion runs at 1,160 to 2,650 tokens per second depending on the compression level.

Those are real numbers, published with the caveat that the Q2_0 row used engine 0.1.36 and the rest used 0.1.26, measured at 4K answers and 32K prompts. The repo also projects that an RTX 3090 with 24 GB of VRAM “should write about 100-140 tokens per second,” which is a projection, not a measurement, and should be read that way.

What is actually new here

Running large Mixture-of-Experts models on consumer hardware is not new. llama.cpp, KTransformers, and a string of forks have been doing it for two years. What Strata packages is the whole stack: an installer that detects the card, picks a quantization that fits the available RAM, downloads roughly 70 GB of weights, and serves an OpenAI-compatible endpoint at http://127.0.0.1:8080/v1 plus an Anthropic-compatible one at /v1/messages. Claude Code, Codex CLI, Cursor, and GitHub Copilot can be pointed at it. There is an MCP server so an AI coding agent can install and start the thing itself.

The mechanism is the interesting part. Qwen3.8-Flash-Next is a sparse MoE model with 24,576 experts, of which roughly 10 activate per token. Strata keeps the most frequently used experts resident in VRAM, holds the full set in system RAM, and lets the CPU work on the rest in parallel. A draft model speculates the next tokens and the full model verifies them in a batch, which the repo says yields a 1.6–1.8x speedup. Long prompts are ingested in chunks of up to 8,192 tokens.

That is a sensible architecture. It is also the same architecture everyone else converged on. The novelty is packaging and defaults, not a breakthrough.

The quantization tax is the story

Read the model-selection table carefully. At 32 GB of RAM the installer recommends “Coder,” a variant with half the experts removed that reaches 91% of the full model’s SWE-bench Verified score, as measured by its authors, and is explicitly weaker on Chinese and other CJK text. At 48 GB you get IQ2_XS or Q2_0 because “the larger sizes do not fit.” Only at 96 GB or more does the installer suggest IQ3_S or Unsloth’s UD-IQ4_XS.

So the headline “125B on a 4090” is really “a 2-bit approximation of a 125B model on a 4090.” That is not a knock. It is the honest framing the repo mostly provides, and the community should adopt it. A 2-bit quantized MoE at 94 tokens/s is genuinely useful for autocomplete, refactors, and chat. It is not the same artifact the Qwen team trained.

The repo is upfront about the worst case: Unsloth’s UD-Q4_K_XL, the closest to full precision, writes only 7–8.5 tokens/s on a 64 GB PC because Strata streams most of it from the SSD. That is a 12x throughput collapse for a quality bump. The tradeoff curve is steep and the repo shows it.

Why this matters for the AI economy

The economics of inference have been consolidating toward the hyperscalers for three years. API pricing from OpenAI, Anthropic, and Google assumes you rent compute by the token. A tool that runs a capable 125B model locally, with nothing leaving the machine, changes the calculus for a specific slice of work: proprietary code, regulated data, offline environments, and anyone whose token bill has crossed into four figures per month.

It also changes the hardware story. Strata’s supported list includes RTX 20-series cards, Radeon RX 6800/6900, and even experimental paths for Tesla P40/V100, Intel Arc, and CPUs without AVX2. That is a long tail of hardware that would otherwise be e-waste. If a 2020-era GPU can serve a 2026-era model, the upgrade cycle for local inference is slower than Nvidia’s roadmap assumes.

What to watch

Three things. First, independent reproduction of the 100–140 tokens/s RTX 3090 projection, which is the number most readers will care about and the one the repo has not measured. Second, whether the “Coder” variant’s 91% SWE-bench Verified figure holds up under third-party evaluation; that number is attributed to the variant’s authors, not to an external benchmark run. Third, whether Qwen, Alibaba’s model team, publishes guidance on running Flash-Next at 2-bit precision, because right now the compression stack is entirely community-built by ISTA-DASLab, UkisAI, and Unsloth.

The repo is MIT-licensed, written by community members, and tested on their own machines. It asks for coffee money. It is not a product with a support contract, and the README says so. Treat the throughput table as a starting point, not a spec sheet, and run your own numbers before you plan a workflow around it.

{/* TODO: verify the 91% SWE-bench Verified figure for the Coder variant against a primary source; currently attributed to “its authors” in the Strata README. /} {/ TODO: verify the RTX 3090 100-140 tokens/s projection; the repo labels it an estimate, not a measurement. */}