senn-techsenn-tech
AI Hardware
AI Hardware2026-10-05· By Franz Senn

Two GB300 towers, two 400G copper cables: what a second box actually multiplies

NVIDIA logo
The GB300 class ships as NVIDIA-licensed DGX Station architecture, built as a tower by ASUS, HP and MSI. StorageReview connected two of these towers for the first time and published throughput data for the pair. (Quelle: Simple Icons, CC0)

StorageReview recently had two GB300-class towers in the lab at the same time: the ASUS ExpertCenter Pro ET900N G3 and the HP ZGX Fury AI Station, two boxes each carrying one NVIDIA GB300 Grace Blackwell Ultra superchip. Brian Beeler linked the pair with two 400G direct-attach copper cables, and what StorageReview describes as the first throughput data for that configuration is now public. We read it closely for two reasons. The mechanism behind the big jumps is the same wall our own 4x RTX 5090 box runs into, just two orders of magnitude higher. And the choice between the two clustering modes in this data is portable across hardware classes.

What is inside the towers, and how they talk

The numbers below come from HP's product page, checked 5 October 2026. Both towers share this silicon, which NVIDIA specifies as the DGX Station architecture.

PathConfiguration
CPUNVIDIA Grace, 72 Neoverse V2 cores, LPDDR5x on four SOCAMM modules
CPU memory496 GB LPDDR5x at 396 GB/s
GPUNVIDIA Blackwell Ultra, 252 GB HBM3e at 7.1 TB/s
Coherent memory748 GB total, 252 GB HBM plus 496 GB LPDDR5x
Computeup to 20 PFLOPS FP4, with FP4 quantisation in the model harness
Network2x QSFP112 at 400 Gb/s each on a ConnectX-8 SuperNIC
Storage2x M.2 PCIe Gen5 off the CPU, 2x M.2 PCIe Gen6 off the ConnectX-8
Power and formdedicated 200/240 V circuit, tower or 5U rails, liquid cooling

The ConnectX-8 is more than a network card: its integrated PCIe switch also feeds the two Gen6 M.2 bays. That looks like a footnote until you see the cabling logic, because disks and network share the same controller.

NVIDIA publishes the pairing as a numbered playbook: cable the rails matching (QSFP0 to QSFP0, QSFP1 to QSFP1), one private /24 per rail at 9000-byte MTU, RoCEv2 point to point, then validate link training, jumbo pings and a real RDMA write test, and finally point NCCL at both rails. StorageReview handed the playbook link to Claude Code, and the agent ran the setup and produced the NCCL report on its own. A practical detail worth noting for evaluation: the instructions are small and precise enough that an agent executes them without error.

NVIDIA's playbook for two stationsTwo 400G DACs, railsmatchedQSFP0 to QSFP0, QSFP1 toQSFP1One private /24 perrailRoCEv2 point to point,9000-byte MTUTraining, ping, RDMAwritemlx5_0 and mlx5_1 mustreport activeNCCL on both railsNCCL_IB_HCA=mlx5_0,mlx5_1
Four steps, driven from a third machine over SSH into both stations. Validation checks link training, jumbo pings and real RDMA writes, not just configuration. (Quelle: NVIDIA playbook Connect two stations, executed by StorageReview)

The fabric itself measured honestly. GPUDirect RDMA from GPU memory to GPU memory hit 392 Gb/s on each of the two rails, 98 percent of line rate, with 1.46 microseconds for an 8-byte write. In NCCL collectives, broadcast and one-way send/receive reached 98 GB/s, all-reduce 92 GB/s and all-to-all 80 GB/s, measured against the roughly 100 GB/s that two QSFP112 cables carry net. For comparison: on the DGX Spark, the ConnectX-7 sat behind two PCIe 5.0 x4 paths and delivered about 200 Gb/s in practice. Here the second tower really does get the full 800 Gb/s onto the table.

Why the second box triples to eleven-times, and never just doubles

The jumps do not come from the faster cable. They come from where the weights live. Take GLM-5.2 in the NVFP4 checkpoint, about 433 GB: on a single tower only 252 GB fits in HBM, so roughly 215 GB of expert weights sit in the LPDDR5x. Every decode step that touches those regions crosses a path with 396 GB/s instead of the HBM's 7.1 TB/s. That is why one tower serves the model at 139 output tokens per second at 32 concurrent requests. Split across two towers in tensor parallel mode, each GPU keeps about half the model in HBM, the spill disappears, and the same concurrency measures 1,525 tokens per second.

Output tokens per second at 32 concurrent requests, short promptsGLM-5.2, 1 tower139GLM-5.2, 2 towers1525GLM-5.3, 1 tower188GLM-5.3, 2 towers853MiniMax-M3, 1 tower1041MiniMax-M3, 2 towers277601600
Identical concurrency, short prompts with 512 input and 512 output tokens. GLM-5.2 and GLM-5.3 spill into the LPDDR5x on a single tower, which is where the multiple comes from; MiniMax-M3 already sits almost fully in HBM, so its gain is plain compute scaling. (Quelle: StorageReview cluster measurement, 5 October 2026)

Two details in the dataset deserve mention for their honesty. First, the headline 27x figure for GLM-5.3 compares one tower at 188 tokens per second with 32 streams against the pair at 5,018 with 256 streams, which is not a like-for-like concurrency comparison. At the same 32 streams the pair measures 853, after 1,301 at 16. Second, every point is a single pass, the curves have dips (GLM-5.3 at 32 streams, MiniMax-M3 at 128), and StorageReview reports them unsmoothed. Tensor parallelism hangs every layer on a collective that must finish before the next step, batching decisions change as concurrency climbs, and you can watch both effects in the data.

Models that fit do not tensor-parallel, they split prefill from decode

For models that fit on one station, tensor parallelism was the weaker choice. StorageReview gave each tower a full copy and separated the phases: the ASUS tower runs prefill and ships the KV cache across, the HP tower decodes without ever pausing for prompt processing. On GLM-5.3 Flash that produced 2,721 instead of 828 tokens per second at 32 streams on the short profile, a 3.3x gain, and 2,448 instead of 1,219 on the prefill-heavy profile, a doubling. DeepSeek V4.1 Flash, with its 188.8 GB of Engram lookup tables and a further 132 GB of expert weights pinned in Grace memory, reached 5,248 output tokens per second at 256 streams on the pair, the highest number in the whole set.

The rule of thumb from the article transfers to any hardware: if the model does not fit into fast memory, split the weights across both GPUs; if it fits, run two full copies and separate prefill from decode. The gain only shows up once there is real demand, at one or two concurrent requests the decode tower works alone at exactly the speed it had before.

What this means for our four RTX 5090s

Our production language lane runs on a host with four RTX 5090 cards and 127 GiB of GPU memory in total. There is no coherent link between CPU and GPU there, the spill path is the PCIe slot, well below half of those 396 GB/s. We run into the same cliff, two orders of magnitude lower: on our hardware quantisation decides whether a model runs at all, while in the GB300 class the number of towers decides whether it runs fast. Buying is not our question to answer, neither HP nor ASUS lists a price on its product page (state of 5 October 2026), and the class serves teams that cannot or will not get a slot in the data centre.

Three carry-aways remain regardless:

  1. The tensor-parallel versus prefill/decode decision rule is hardware-independent, and we can check it against our own lane curves as soon as two GPU hosts each hold a full copy.
  2. The number that decides throughput once weights spill is the bandwidth of the spill path, not the petaflops on the datasheet. 7.1 TB/s against 396 GB/s is an 18x cliff, and that cliff is exactly what turns the second tower into a multiple.
  3. The fabric playbook (copper instead of optics, RoCEv2 point to point, MTU 9000, one /24 per rail, validation with ib_write_bw) is documented and agent-executable. But tensor parallelism across machine borders needs this fabric in the 100 GB/s class; that is a dedicated fabric, not an add-on to a general-purpose network.

Further reading

Questions?
How fast is the link between two GB300 stations in practice?+

StorageReview measured 392 Gb/s per rail with GPUDirect RDMA, GPU memory to GPU memory, which is 98 percent of the 400G rating, and 1.46 microseconds for an 8-byte RDMA write. In NCCL collectives, broadcast and one-way send/receive reached 98 GB/s against roughly 100 GB/s that two QSFP112 cables carry net. The run dates from 5 October 2026.

Does a second station always bring a multiple?+

No. Models that overflow the GPU memory and spill into DRAM gained 11x at 32 concurrent requests (GLM-5.2). Models that already fit on one station gained 2 to 3.3x when prefill and decode were split across the pair, and roughly nothing at one or two concurrent requests.

Do these results carry over to our hardware?+

The mechanism carries over, the numbers do not. Our RTX 5090 host has 127 GiB of GPU memory in total and no coherent CPU-GPU link, so its spill path is the PCIe slot. On our hardware quantisation decides whether a model runs at all; in the GB300 class the number of towers decides whether it runs fast.