KV cache in 4 bits: 8.8 instead of 4.7 full sessions on the same four RTX 5090s
Since the evening of 2 October our production lane for Qwen3.8-Flash-Next runs with a 4-bit KV cache. The pool grows from 1,041,901 to 1,944,081 tokens. At 220,000 tokens of context per request that means 8.8 instead of 4.7 full sessions fit into GPU memory at the same time. Answer quality stayed unchanged in every check, single-stream speed dropped by eight percent. Same model, same four RTX 5090s, same memory budget per card.
What we changed
The KV cache is the part of GPU memory where the model keeps the history of a request. Until now we stored it in FP8, one byte per value. The int4_per_token_head switch stores every value in 4 bits and keeps the scaling factors per token and attention head. That works out to 264 bytes per head and token at a head dimension of 256, roughly half of the previous footprint.
Two details had to be solved for this. Our vLLM build already knows the write path for INT4, but the model's sparse-attention layers kept reading the cache as FP8. We brought the read side to the same level and verified it against the reference implementation. On top of that, the server refused to start with the usual block size of 800 tokens. The Mamba state takes 408,576 bytes per session and shares pages with the KV cache. An 800-token page holds only 211,200 bytes at INT4. With a block size of 1600 a page fits 422,400 bytes, and the server starts.
The same Mamba state is also why the pool grows by a factor of 1.7 to 1.87 instead of exactly two. Its footprint stays constant while the KV part shrinks.
The measured values
Decoding is slower with INT4 in a single stream: 90.7 against 98.6 tokens per second, measured on the same prompts at temperature zero. Under parallel load this hardly matters, because requests wait less often for free cache space. The first start with the new image took almost ten minutes, because new kernel paths are compiled. Later starts take about five.
The number of concurrent requests stays at eight, that cap is unchanged. What changes: all eight slots may now hold the full context at the same time. Before, eight slots had to share 4.7 full sessions, and very long documents had to wait or be shortened.
How we checked quality
Four checks, all run on 2 October. A kernel harness compares the INT4 output directly against the torch reference: minimum cosine similarity 0.99982. Then needle tests, where a hidden sentence has to be recovered from long texts. The needle was found at 180,000 tokens in the middle and at 215,000 tokens at depths 0.3 and 0.9. Deeper is not possible on this configuration, because the engine caps at 220,000 tokens.
Four coding tasks at temperature zero returned the same algorithms and structures on both cache formats, with small differences in variable names and comments. We extracted and ran two of the solutions, an LRU table and a Dijkstra, both functionally correct. The last test was a mixed run with four parallel requests of about 128,000 tokens each, which completed cleanly.
The kernel incident on the same day
The day had a second construction site. A routine run of apt full-upgrade installed kernel 7.0.0-38. In the process DKMS rebuilt the NVIDIA module, from the unmodified source tree. The patch that enables direct memory transfer between our four cards was missing afterwards. Without it, traffic between the cards goes through host memory, and that direct path is exactly what makes tensor parallelism fast.
We built the patched modules from our clone of open-gpu-kernel-modules against the new kernel headers, backed up the original modules first, rebooted and checked with nvidia-smi topo -p2p r: all card pairs OK again. After that we pinned the kernel and NVIDIA packages with apt-mark hold. A kernel change now happens only deliberately, with the rebuild in the same maintenance window.
What this means for operations
The bottleneck of the last weeks was memory for long parallel requests, not compute. With almost nine full sessions in the cache, long documents land in the queue less often, and the agents that read and write for us all day get their complete context. The FP8 image stays on the machine, along with a snapshot of the state before the switch. If quality issues show up with very long contexts in daily use, the lane goes back to FP8 with one command until the cause is found.
Further reading
- vLLM on GitHub
- aikitoria/open-gpu-kernel-modules: the P2P patch for GeForce cards
- AI with Eric: Qwen3.8-Flash-Next sparse attention under test (video, 27 Aug 2026)
- Cloud Codes: Qwen3.8-Flash-Next and the VRAM limit (video, 28 Aug 2026)
- The build comparison with FP8 KV from 24 September
- How FP8 KV reached our production
Why does the pool not exactly double?+
The linear part of the model, the Mamba state, takes 408,576 bytes per session and shares memory pages with the KV cache. Its footprint stays the same while the KV part is halved. Measured across boots we land between 1.7 and 1.87 times the old pool.
Does answer quality suffer with a 4-bit KV cache?+
Not in our checks from 2 October. The kernel matched the reference with a cosine similarity of 0.99982, needles hidden in 215,000-token texts were found, and four coding tasks came back functionally identical. We still watch production, because no test covers everyday use completely.
What happens if a problem shows up in production anyway?+
The way back is a single command and takes a good six minutes. The FP8 image and a snapshot from 2 October are kept on the machine for exactly that case. The lane then runs again with the 1,041,901-token pool.
senn-tech