senn-techsenn-tech
KI-Hardware
KI-Hardware2026-10-02· By Franz Senn

KV cache in 4 bits: 8.8 instead of 4.7 full sessions on the same four RTX 5090s

Since the evening of 2 October our production lane for Qwen3.8-Flash-Next runs with a 4-bit KV cache. The pool grows from 1,041,901 to 1,944,081 tokens. At 220,000 tokens of context per request that means 8.8 instead of 4.7 full sessions fit into GPU memory at the same time. Answer quality stayed unchanged in every check, single-stream speed dropped by eight percent. Same model, same four RTX 5090s, same memory budget per card.

Qwen logo
Four bits per value instead of eight: the Qwen3.8-Flash-Next KV cache halves its footprint, the Mamba state stays as it is. (Quelle: Alibaba Cloud / Wikimedia, CC0)

What we changed

The KV cache is the part of GPU memory where the model keeps the history of a request. Until now we stored it in FP8, one byte per value. The int4_per_token_head switch stores every value in 4 bits and keeps the scaling factors per token and attention head. That works out to 264 bytes per head and token at a head dimension of 256, roughly half of the previous footprint.

Two details had to be solved for this. Our vLLM build already knows the write path for INT4, but the model's sparse-attention layers kept reading the cache as FP8. We brought the read side to the same level and verified it against the reference implementation. On top of that, the server refused to start with the usual block size of 800 tokens. The Mamba state takes 408,576 bytes per session and shares pages with the KV cache. An 800-token page holds only 211,200 bytes at INT4. With a block size of 1600 a page fits 422,400 bytes, and the server starts.

The same Mamba state is also why the pool grows by a factor of 1.7 to 1.87 instead of exactly two. Its footprint stays constant while the KV part shrinks.

The measured values

KV pool in tokens at the same byte budget per cardFP8 (September state)1041901 · 4.74 sessions at 220kINT4 (since 02 Oct)1944081 · 8.84 sessions at 220k02000000
Server startup messages, both runs with an identical cache budget and 220,000 tokens of context. (Quelle: senn-tech, own measurement from 2 October 2026)
Single-stream decode, tokens per secondFP898.6 · 98.6 tok/sINT490.7 · 90.7 tok/s (−8%)0110
Same prompts at temperature zero, both runs on 2 October 2026. (Quelle: senn-tech, own measurement from 2 October 2026)

Decoding is slower with INT4 in a single stream: 90.7 against 98.6 tokens per second, measured on the same prompts at temperature zero. Under parallel load this hardly matters, because requests wait less often for free cache space. The first start with the new image took almost ten minutes, because new kernel paths are compiled. Later starts take about five.

Cloud Codes from 28 August 2026 on the model's memory math: why on consumer cards the cache, not the weights, is the bottleneck.

The number of concurrent requests stays at eight, that cap is unchanged. What changes: all eight slots may now hold the full context at the same time. Before, eight slots had to share 4.7 full sessions, and very long documents had to wait or be shortened.

How we checked quality

Four checks, all run on 2 October. A kernel harness compares the INT4 output directly against the torch reference: minimum cosine similarity 0.99982. Then needle tests, where a hidden sentence has to be recovered from long texts. The needle was found at 180,000 tokens in the middle and at 215,000 tokens at depths 0.3 and 0.9. Deeper is not possible on this configuration, because the engine caps at 220,000 tokens.

Four coding tasks at temperature zero returned the same algorithms and structures on both cache formats, with small differences in variable names and comments. We extracted and ran two of the solutions, an LRU table and a Dijkstra, both functionally correct. The last test was a mixed run with four parallel requests of about 128,000 tokens each, which completed cleanly.

External test from 27 August 2026: it shows when the model's sparse attention loses content from long contexts, and which setting catches that. Our needle checks measure exactly this risk.

The kernel incident on the same day

The kernel incident on 2 Octoberapt full-upgradekernel 7.0.0-38DKMS rebuildP2P patch missingRebuild the modulesagainst new headersReboot and checkP2P OK on all pairsapt-mark holdpackages pinned
The rebuild took a quarter of an hour, after which the lane ran on the fast path again. (Quelle: senn-tech, own protocol from 2 October 2026)

The day had a second construction site. A routine run of apt full-upgrade installed kernel 7.0.0-38. In the process DKMS rebuilt the NVIDIA module, from the unmodified source tree. The patch that enables direct memory transfer between our four cards was missing afterwards. Without it, traffic between the cards goes through host memory, and that direct path is exactly what makes tensor parallelism fast.

NVIDIA logo
The P2P patch lives outside the NVIDIA source tree. Every kernel update rebuilds the module, and without the patch the direct path between the cards is closed again. (Quelle: Simple Icons, file nvidia.svg (CC0))

We built the patched modules from our clone of open-gpu-kernel-modules against the new kernel headers, backed up the original modules first, rebooted and checked with nvidia-smi topo -p2p r: all card pairs OK again. After that we pinned the kernel and NVIDIA packages with apt-mark hold. A kernel change now happens only deliberately, with the rebuild in the same maintenance window.

What this means for operations

The bottleneck of the last weeks was memory for long parallel requests, not compute. With almost nine full sessions in the cache, long documents land in the queue less often, and the agents that read and write for us all day get their complete context. The FP8 image stays on the machine, along with a snapshot of the state before the switch. If quality issues show up with very long contexts in daily use, the lane goes back to FP8 with one command until the cause is found.

Further reading

Questions?
Why does the pool not exactly double?+

The linear part of the model, the Mamba state, takes 408,576 bytes per session and shares memory pages with the KV cache. Its footprint stays the same while the KV part is halved. Measured across boots we land between 1.7 and 1.87 times the old pool.

Does answer quality suffer with a 4-bit KV cache?+

Not in our checks from 2 October. The kernel matched the reference with a cosine similarity of 0.99982, needles hidden in 215,000-token texts were found, and four coding tasks came back functionally identical. We still watch production, because no test covers everyday use completely.

What happens if a problem shows up in production anyway?+

The way back is a single command and takes a good six minutes. The FP8 image and a snapshot from 2 October are kept on the machine for exactly that case. The lane then runs again with the 1,041,901-token pool.