senn-techsenn-tech
AI Hardware
AI Hardware2026-10-03· By Franz Senn

Two DGX Spark boxes on our bench: DeepSeek V4-Flash runs, V4.1-Flash does not fit

At the end of September we had two NVIDIA DGX Spark boxes on the bench, on loan, cabled together through their QSFP ports. The question we wanted answered was narrow: can a pair of these cover one of our open cases, meaning a large open-weights model in our own machine room, and what would it cost us. On 28 September 2026 we served two DeepSeek models from them with vLLM. DeepSeek-V4-Flash ran at a measured 63.7 tokens per second on a single stream. DeepSeek-V4.1-Flash failed on memory, and the reason behind that failure applies to every box with unified memory.

Checkpoint size and memory in GB4× DGX Spark, total512 · 512DeepSeek-V4.1-Flash, checkpoint510.3 · 510.33× DGX Spark, total384 · 3842× DGX Spark, total256 · 2562× DGX Spark, usable, measured186 · 186DeepSeek-V4-Flash, checkpoint159.6 · 159.60560
Four boxes would total 512 GB, and the checkpoint alone is 510.3 GB. That leaves nothing for the operating system, the KV cache or activations. The 186 GB figure is our own reading from the two-box setup. (Quelle: NVIDIA DGX Spark specifications, Hugging Face, our own measurement 28 Sep 2026)

What the spec sheet counts, and what it leaves out

NVIDIA quotes parameters. The product page says 128 GB and up to 200 billion parameters for one box, 256 GB and up to 400 billion for two, 512 GB and up to 700 billion for four. ConnectX-7 links up to four of them. That ladder is a buying aid. What decides a purchase is a different quantity: how many bytes one parameter of this specific model occupies in memory.

ModelParametersCheckpoint on diskBytes per parameter
DeepSeek-V4-Flash284 B159.6 GB, 46 safetensors0.56
DeepSeek-V4.1-Flash763 B510.3 GB, 48 safetensors0.67

Both figures come straight from DeepSeek's Hugging Face repositories as file sizes of the safetensors shards. At 0.67 bytes per parameter, the 400 billion parameters NVIDIA advertises work out to 268 GB, which is past what two boxes hold. Between one model generation and the next, the step in parameters is also a step in byte arithmetic, and the byte arithmetic is the one that decides the purchase.

A 64 GB model arrived on 2 October

On 2 October NVIDIA announced a second configuration: 64 GB of LPDDR5X at 4,999 US dollars, sold only through Acer, ASUS, Dell, Gigabyte, HP and MSI, in shops from 23 October. Nothing changes in the compute or the networking, it is still a GB10 at 273 GB/s with ConnectX-7 attached. One box is meant to carry models up to 100 billion parameters, two clustered boxes pool to 128 GB and are quoted up to 200 billion.

That makes the new SKU a clean test of the unit problem. Two 64 GB boxes hold 128 GB together, our measured usable range was 186 GB, and the DeepSeek-V4-Flash checkpoint is 159.6 GB. The cheap pair therefore misses precisely the model we measured on two large boxes on 28 September, even though on paper it reaches the same memory figure and the same parameter tier. Per byte it is also the dearer buy: 4,999 dollars across 64 GB is 78 dollars per GB, where the 128 GB model at 6,950 dollars works out at 54 dollars per GB.

The arithmetic behind 63.7 tokens per second

When a model generates text, every weight a token needs passes through memory once. DeepSeek-V4-Flash activates 13 billion parameters, and at 0.56 bytes that is 7.3 GB per token. One DGX Spark moves 273 GB/s across its LPDDR5X pool, so two boxes reach 546 GB/s in theory. Divided out, that is 74.7 tokens per second as the ceiling, and we measured 63.7, which is 85 percent of it. A single box would sit at 37.4 tokens per second.

Landing that close to the bandwidth ceiling fits what we found in our comparison of the desktop boxes: this class of machine is a memory-bandwidth device that happens to have a GPU attached, not a compute device. Its one petaFLOP of FP4 throughput spends most of the decode path waiting. The same arithmetic for our own hardware, including how much of the rated bandwidth actually arrives, is in our RTX 5090 memory bandwidth measurement.

The figure from the October announcement fits this picture. NVIDIA puts the gain of two clustered boxes over a single one at up to a factor of 1.7 on Qwen3.8 27B, explained by the doubled memory bandwidth. For DeepSeek-V4-Flash, our measured 63.7 tokens per second sit against 37.4 tokens per second for one box on paper, which by definition is 1.7 as well.

Why V4.1-Flash does not fit

DeepSeek-V4.1-Flash has 552 billion backbone parameters. On top of those sit 196 billion Engram parameters, and Hugging Face counts the sum as 763 billion. Engram is a lookup memory. The last few tokens get hashed, the hash values index large static tables, and the rows that come back are mixed into the hidden state through a gate. The model never computes over those parameters, it reads them, which is how capacity grows without adding cost per token.

For memory planning the consequence is unpleasant. The vLLM documentation describes this model's tables as two Engram layers of roughly 384 million rows each, with FP8 entries of 256 bytes plus their block scales, about 200 GB of table weights on their own. That is 39 percent of the checkpoint. The same documentation states that the layer layout, the table sizes and the n-gram size come from the checkpoint and are not operator-configurable. There is no setting that makes the tables smaller.

Put that against our setup: 510.3 GB of checkpoint against 256 GB of shared memory, 186 GB of it usable. Even if the backbone weights are split across both boxes, the lookup half has to stay resident somewhere. Four boxes at 512 GB would be the next step, and NVIDIA does support four, but only through a switch. With 510.3 GB of weights, nothing is left for the operating system, the KV cache or the activations. DeepSeek-V4.1-Flash is simply not a model for this hardware class.

The offload trick does not exist on this hardware

vLLM moves the Engram tables into pinned host memory by default, looks them up directly from the Triton kernel through unified virtual addressing, and prefetches on a second CUDA stream. The logic is sound on a machine with graphics cards: the expensive memory is the memory on the cards, and moving the tables out frees it for the KV cache. In our lane with four RTX 5090 that offload is real relief, because two separate memory worlds sit side by side there, with very different bandwidth and a very different price per byte.

The DGX Spark does not have those two worlds. NVIDIA describes its 128 GB of LPDDR5X as coherent unified system memory, and it is exactly that: tables, weights, KV cache and activations live on the same chips and share the same 273 GB/s. Offloading here moves an address range into another address range. It does not release memory, and the lookups then take bandwidth from the same pool the weights are already using. A technique that saves 200 GB on one class of hardware is worth nothing on the other. Anyone sizing one of the new desktop boxes against a current-generation DeepSeek model has to carry that through the calculation.

Tables on the SSD and a 4-bit file, worked through

On 3 October we counted the files of the checkpoint through the Hugging Face API, because we did not want to keep guessing at the 200 GB. There are 96,085 tensors in 48 files, 510.3 GB in total. The last two files are 101.5 GB each, and the index file names what is inside them: layers.1.engram.embed.weight and layers.14.engram.embed.weight together with their scales. Those are the tables, 203.0 GB, 39.8 per cent of the file. The remaining 307.3 GB are the weight network, the vision tower and the embeddings.

That turns the plan of putting the tables on an SSD into arithmetic with known numbers. 307.3 GB would stay in memory. Three boxes hold 384 GB on paper, while our measurement on 28 September pointed closer to 93 GB per box, which is 279 GB across three. The weight network on its own fills those 279 GB, with nothing left for the KV cache or the operating system.

The second lever, quantisation, has already been pulled where the file is widest. The configuration of the original carries expert_dtype fp4, so the experts are computed in 4 bits already, and that is what makes 0.67 bytes per parameter add up. There is no second 4-bit step available for those same experts. We measured a quantised build, W4A16-Engram-AutoRound from 14 September 2026: 451.7 GB against 510.3 GB. The tables halve to 111.2 GB, while the other 46 files grow from 7.4 to 8.2 GB each because W4A16 keeps the activations at 16 bits. Under the line that is 58.6 GB less, 11.5 per cent.

The SSD is the tighter ceiling, ahead of capacity. A table entry is 256 bytes and an NVMe drive answers in blocks of 4 kilobytes, so every lookup drags along sixteen times as many bytes. Even 200,000 random reads per second, a figure we have not measured on the built-in drive, would give 0.8 GB/s against 273 GB/s inside the box. The measurement worth running is a single box offloading to its own drive, about half a day.

The second ceiling: the link between the boxes

Each QSFP port delivers 200 Gbit/s, which is 25 GB/s. Two cables between two boxes make 50 GB/s, measured against 546 GB/s inside the two memory pools, a factor of eleven. Tensor parallelism across boxes pushes activations over that link on every layer, which is why the second box is not twice the speed but first of all extra memory. The 74.7 tokens per second ceiling only holds if the distribution works perfectly. Our first start across the two boxes came up with 2 of 8 ranks and ended in a SIGFPE inside NCCL. The second attempt, with the network configuration adjusted, produced the 63.7 measurement.

From the third box onward the cabling itself becomes a question. Every box has exactly two QSFP ports, which is why NVIDIA describes connecting three boxes as a ring with three cables. Each box sits against both neighbours, and with three nodes a ring is at the same time the complete connection, every box directly attached to every other. In a chain of two cables the two end boxes would have no link between them, and there is no device in between to forward anything. Direct cabling ends at three boxes, four boxes only work across a switch. The ring triples the memory, the attachment per box stays at two ports and therefore at 50 GB/s.

The cost question is the same arithmetic. The pair came to 11,258 euros at street price in early September, 5,629 euros per box, laid out in our hardware comparison from 6 September. List prices have moved since: the 128 GB Founders Edition has stood at 6,950 US dollars since 2 October, after 3,999 at launch and 4,699 in the spring, and German shops listed 5,799 euros on the same day. The small model is what makes pairing worth thinking about. Two 64 GB boxes cost 9,998 dollars and deliver the same 128 GB of memory and the same 546 GB/s as two large ones, while one large box costs 6,950 dollars and brings 273 GB/s. Anyone buying bandwidth who can live with 128 GB buys two small ones.

For the money of the large pair our own lane runs several graphics cards on one mainboard, with separate card memory and system memory, which is where the offload lever described above exists at all.

Where V4.1-Flash really does move ahead

The KV cache. DeepSeek states 890 bytes per token for this model's global KV cache, one quarter of DeepSeek-V4-Flash and around one part in 437 of DeepSeek-V1. Across a one million token context that is 0.93 GB, where V4-Flash at four times the rate would need 3.7 GB. For the case we actually care about, agents with long contexts on small hardware, that is the more interesting number on the datasheet. Those two quantities, decode throughput and memory per token, were also the subject of our comparison of two vLLM builds with a quantised KV cache on our own lane.

What we take from this

DeepSeek-V4-Flash in its current weight mix is a candidate for our own lane, and 159.6 GB of weights is a number one can plan with. We will run the serving setup on the four RTX 5090 and measure decode throughput and concurrent operation.

For the DGX Spark class the recommendation stands as before. One box is a throughput path for a single model whose checkpoint stays well under half the memory, plus the KV cache for the contexts we actually run. A second box doubles the memory and doubles the bandwidth, which turns into up to a factor of 1.7 in decode. What it does not add is a second memory pool.

One rule goes onto our evaluation checklist and applies everywhere from here: we calculate with checkpoint bytes and measured bytes per parameter, not with the parameter figure from a press release. Between DeepSeek-V4-Flash and V4.1-Flash the parameter count grows by a factor of 2.7 and the bytes on the wire by a factor of 3.2. Still open on our side is how a pair of DGX Spark boxes behaves under concurrency, tokens per second per request at four and eight simultaneous sessions. Our run on 28 September did not produce that figure cleanly.

Sources

Questions?
Which DeepSeek model runs on two DGX Spark boxes?+

DeepSeek-V4-Flash, 284 billion parameters, 159.6 GB of weights on disk. On 28 September 2026 we measured 63.7 tokens per second on a single stream, against a theoretical ceiling of 74.7 tokens per second set by the memory bandwidth of the two boxes.

Why will DeepSeek-V4.1-Flash not fit on two DGX Spark boxes?+

Its checkpoint is 510.3 GB and the two boxes share 256 GB of memory, of which 186 GB were usable in our setup. Roughly 200 GB of that is taken by the model's Engram lookup tables, whose size is fixed inside the checkpoint and cannot be trimmed by the operator.

Is a DGX Spark worth buying for us?+

As a single box, for one model whose checkpoint stays well under half the memory, plus the KV cache the contexts we actually run need. Two boxes double memory and bandwidth alike, but traffic across the link between them runs at 50 GB/s against 546 GB/s inside the two pools. The cheapest way to get that pair is two 64 GB boxes, as long as the model fits in 128 GB. For our requirement, an x86 lane with several graphics cards keeps more room to move.