senn-techsenn-tech
AI & Development
AI & Development2026-10-11· By Franz Senn

Strata runs a 125-billion model at 50 to 60 tokens per second on a graphics card from 2020

A self-test is going around this week, and the numbers sound too good: a PC with an AMD RX 6800, 16 GB of video memory, 64 GB of RAM, a graphics card that shipped in 2020. In August it ended at 3 to 4 tokens per second under Ollama, with an agent nobody would use. This week the same machine writes 50 to 60 tokens per second, and the agent works. In between sits two weeks of software, a backend called Strata. I read both of the author's posts, measured the project's current state, and pulled up our own measurements of the same model family. The self-test is honestly written. The headline only partly survives that. A second self-test has appeared since: same backend, different hands, a cheaper class of machine.

Two different models sit between August and today

The author started in August with qwen2.5-coder:14b (a 32K context window, while his agent demands at least 64K), moved to gpt-oss:20b and qwen3.5:9b-q8_0, which threw his questions back at him, then ran gemma4:31b down three quantization steps to a 13 GB file while Ollama kept 19 to 28 percent of the work on the CPU and the box went unstable. A smaller mixture-of-experts model at least answered in seconds instead of minutes, but could not get its tool calls to the CRM right. His August verdict: speed and quality trade against each other one for one, and at 3 to 4 tokens per second you fall asleep at the keyboard.

The detail that decides how you read the headline: today's model, Qwen3.8-Flash-Next, does not run on Ollama at all. The request to support this architecture is open at llama.cpp (issue 27741) and at Ollama (issue 18071), both checked on October 11, 2026. So "3 to 4 versus 50 to 60" compares two different models on two different engines. That does not make the end value wrong. It does mean the gain belongs to the new backend and the new model together.

How Strata spreads the weights across GPU, RAM and SSD

Strata (GitHub Niko1221/Strata, MIT license) is a C++ engine with its own kernels on top of a vendored copy of the ggml library, not a llama.cpp fork. State measured through the GitHub API on October 11, 2026: the repo was created September 24, the newest release v0.1.43 came the same day, 21,456 stars, 2,020 forks, 214 named contributors. Open threads stand at 687, and that is 333 issues plus 354 merge requests. The counter most dashboards show as an issue count adds the two together. The engine serves exactly one model family, and that family was built for offloading from the start.

The model card gives the numbers: 125 billion parameters, 6 billion active per token, 512 experts per layer, 24,576 experts counted across all layers. On top of that a 51-billion n-gram embedding field, which the card itself describes as a better offloading candidate than the MoE layers, a 4-billion head for speculative decoding, and a native context of 262,144 tokens.

Where Strata keeps the weightsGraphics cardattention, routers, sharedexperts, part of the KVcache, and a cache of theSystem RAMall 24,576 experts, pinned,so the card never waits onthe diskProcessorcomputes the experts thatare not on the card, at thesame time as the card worksSSDa lookup table; a tokenreads only a few small rowsfrom it
Per the project's own description in docs/HOW_IT_WORKS.md. Every extra gigabyte of VRAM holds roughly 700 more experts in the cache. (Quelle: Strata, docs/HOW_IT_WORKS.md)

That this axis is the real one I measured myself in our four-day test: two more RAM sticks, from 62 to 94 GiB, lifted decode speed by 18 percent, because this model reads its n-gram table from system RAM for every single token. Strata turns that lever into its whole design.

What the author tuned, and what each step was worth

According to his write-up it was four measures, in this order:

  1. --vram-reserve-mib 1500: gives the desktop 1.5 of the 16 gigabytes back. The workstation's freezing stops, and throughput moves from 30 to 40 up to 40 to 50 tokens per second.
  2. The PCIe slot: the card sat in the lower slot, physically x16, electrically only x4 through the chipset. After the move to the slot wired directly to the CPU, he measured roughly 12 GB/s of bus bandwidth instead of roughly 3, and the rate settled at a steady 50 to 60. It is the single biggest jump in his report, and at heart a cable problem that took him the long way to find.
  3. --draft-vocab en: the draft model keeps English vocabulary only, which frees room for more cached experts. He puts the gain at about one percent and says himself it is not measurable.
  4. Context from 64K to 128K: no loss he can see. The expected loss of around 20 percent at 256K is his estimate, and he says so himself; he did not measure it. For his agent (Hermes) he set fix_max_tokens to true, because it asks for the full context size with every call.

The field test after the tuning: his agent parsed ODS spreadsheets, kicked off web searches from what it read in them, and summarized the results. That is exactly what the smaller models had failed at in August. Every throughput number in this chain is self-reported; the author publishes neither raw data nor screenshots, which is worth stating fairly either way.

The project's own tables put it in range

The README lists speeds per model size in two tables, one per vendor, at 4K answers over 32K prompts. His size, IQ2_XS: an NVIDIA RTX 5070 (12 GB) at 79 tokens per second, an AMD RX 9070 XT (16 GB) at 52. The author's old RX 6800 beats the vendor figure for the current AMD card. His machine is not a driver path coaxed past its limits, it is exactly the hardware class the project was written for (the supported list in the questions below names the RX 6800 series outright).

A second build: two graphics cards for 984 dollars

On October 10, the night after my measurement run, the channel Digital Spaceport showed the same backend on a second-hand Dell Precision T5810 with two RTX 3060 cards of 12 GB each and 64 GB of RAM. His parts list sums to 984 dollars, processor included at 7. Twelve gigabytes is the floor of the project's install list. The channel has run since 2008, names 103,000 subscribers and discloses its Amazon and eBay commissions. In his GPU tracker the RTX 3060 is the only card listed with exactly 12 GB, and the three cheaper entries there, P100, P40 and V100, are Tesla cards Strata only admits through its experimental CUDA 12 engine.

He ran all four sizes against one prompt, an animated graphic, one generation each. The repository documents a run on a single RTX 3060, Core i7-12700, DDR4-3200, Strata v0.1.41:

Sizeone 3060, run in the repotwo 3060s, his video
Q2_042.963.3
IQ2_XS45.562
IQ3_XXS26.252.4
IQ3_S20.2not measured

Tokens per second. His memory is slower, his processor dates from 2016, and he still lands above the single card. The procedure in docs/MULTI_GPU.md predicts that direction: every card keeps an expert cache for its own layers, in his words "pipeline (layer) parallelism, not tensor parallelism", no NVLink involved. The comparison is not controlled: different hosts, different engine builds. It supports the direction while the factor stays open. The same document shows it holds on cheap parts: two Tesla P40s lifted IQ2_XS decode from 19.6 to 34.8 tokens per second, one measured report.

His methodology page says in so many words that his tests are "by no means scientific", presented as "so I do it like this". Of the largest size he reports 53.4 GB of memory and no throughput, and he measures speed, he writes, without a refined method.

The limits are documented, candidly

  • One request at a time: by default Strata answers one request while the others wait. The optional parallel: 2 setting makes every single answer slower on a 12 GB card. The author of the single-card test describes the practical consequence himself: a small summarization call from his agent queues up ahead of his actual work.
  • Every restart drops the expert cache, and the first large request starts from zero again. Budget about one minute for the first message of a conversation per 30,000 tokens.
  • No external security audit has happened so far, and the repository says so itself. 333 open issues and 354 open pull requests on a two-week-old repository are the same finding from the other side (correction, October 11: this line previously carried one combined figure of 690 open items, which mixed the two).
  • Two bug reports match our own traffic shape exactly and were open on October 11: #606 wedges the engine after one large generation at 155K context, so that every later request returns a single repeated character while the health endpoint stays green; #481 is a deadlock mid-stream that the stall watchdog never catches. The defect we disliked most, #528, which charged four to six times the decode time on continued conversations, landed exactly on our sustained-dialogue load, and was closed on October 7.
  • All advertised sizes sit at 2 to 3 bits (Q2_0, IQ2_XS, IQ3_*). The "Coder" variant keeps 256 of the 512 experts per layer; its quality promise (91 percent of the full model on SWE-bench Verified) comes from the authors themselves and is unconfirmed, and the README notes the weakness outside English. Nobody has measured any of these sizes against German-language agent or mail traffic, us included.

The same model family in our own rack

We have run Qwen3.8-Flash-Next in-house since late August, and the measurements are published on this blog: a single request under llama.cpp came in at 96.8 tokens per second, eight concurrent ones managed 50.7 together (four-day test). On vLLM with expert parallelism the same model delivers 205 tokens per second across 32 concurrent streams (the KV-cache patch post). Strata against that is no contest, because they solve different jobs: 50 to 60 tokens per second at one desk that nobody queues behind, versus 205 across two dozen agents at once in our own AI stack. Strata does not bundle, and bundling is precisely where our load lives.

Our assessment from October 4, documented and public here for the first time: verdict WATCH, nothing installed. One of the reasons was wrong, and I withdraw it (correction at this point, October 11). We had written that the fast CPU kernels demand AVX-512, which no free machine of ours has. The install list asks for AVX2 and calls AVX-512 "a bit faster"; one size has a fast path with it. The builds they ship are AVX2, and two community runs in the repository sit on Xeon chips without AVX-512, one on exactly the chip of the build above. Our processor was never the obstacle. Two further reasons stand: multi-GPU means splitting layers and experts, not tensor parallelism. And #606 plus #481 hit the traffic signature of our agents (long prompts, tool calls inside live streams, deliberately canceled requests). That is a fit verdict on hardware and topology, not a quality judgment about the authors, whose openness I want to call out as genuinely good.

Who can use this today

For single seats whose data should not leave the building, the self-test turned into a usable machine: a 2020 graphics card, 64 GB of RAM, 80 GB of SSD, an agent that reads spreadsheets, searches and summarizes, at reading speed. Anyone reproducing it should budget two things that no tokens-per-second figure shows: a slot actually wired to the CPU, and the first request after every restart. For multi-seat duty, reliability and auditable paths, the requirements from our family portrait still stand, and Strata has not met any of them in two weeks. The architecture question stays open and on our watchlist: the unanswered request at llama.cpp shows the single-model trick is still up for adoption there. Once the architecture lands upstream, the same models start their race under the same engines from the front line again, and we will report with fresh measured numbers.

Further reading

Questions?
Is "3-4 tokens per second becomes 50-60 on the same machine" a fair Ollama-versus-Strata comparison?+

No. In August, Ollama ran a different model, a 31-billion Gemma in several quantizations. Today's run uses Qwen3.8-Flash-Next on Strata. The request to support that architecture is still open at both Ollama and llama.cpp, so Ollama never got to compete at equal terms. The 50 to 60 tokens per second remain a self-reported figure, but a carefully documented one.

What hardware do you need to reproduce it?+

The project's own install list: 12 GB of VRAM or more (NVIDIA RTX 20 through 50 series; AMD RX 6800 and 6900 series, RX 7800 XT and 7900 series, RX 9060 XT and 9070, Radeon AI PRO R9700), at least 32 GB of system RAM, roughly 80 GB of disk, Windows or Linux. The author's two biggest gains cost nothing: a graphics-card slot actually wired to the CPU, and a VRAM reservation for the rest of the desktop. Two cards need no NVLink: the engine splits by layer, each card serving its own block, and cards sitting in x4 or x1 slots work as well.

Could a company replace ten cloud seats with this?+

That is not what it was built for. Strata answers one request at a time by default, and its optional parallel mode makes each individual answer slower on a 12 GB card. The project states it has had no external security audit, and one open bug wedges the engine with the health endpoint still reporting green. For a single desk whose data should not leave the building, it has become usable. The model card itself recommends vLLM, SGLang or KTransformers for throughput duty.