Four DGX Sparks, No Switch: GLM‑5.3‑Flash NVFP4 on a Ring

Four NVIDIA DGX Sparks, four direct ConnectX‑7 cables, no Ethernet or InfiniBand switch in the data path, and GLM‑5.3‑Flash NVFP4 serving at 68.6–70.3 tokens/s on prose, 126.6–138.8 tokens/s on code, and 4,499 tokens/s on a 21.9K-token cold prefill.

That is the result I reached on October 10, 2026. The interesting part is not just that the four-node ring works. It is that it is faster than the same model on my three-node full mesh, uses the stock NCCL 2.30.7 build, keeps quality intact, and exposes a 4.24M-token KV pool, without buying a switch.

The repository, launchers, fabric setup, tests, and benchmark battery are in calvarado2004/sparks.

Short version. Moving from three Sparks in a triangle to four Sparks in a switchless ring improved single-stream code decode by 34–47%, prose by 40–43%, 21.9K-token prefill by 53%, and doubled the useful KV pool from 1.85M to 4.24M tokens. A real 54,214-token one-shot HTML generation ran at 80.4 tok/s.

Four DGX Sparks wired as a physical ring. Each Spark has one direct cable to each neighbor.
Four DGX Sparks wired as a physical ring. Each Spark has one direct cable to each neighbor.

Why try a ring at all?

My previous GLM‑5.3‑Flash card used three Sparks in a full mesh. With three nodes, every rank can have a direct cable to every other rank, which is a comfortable topology for tensor parallelism. Adding a fourth node changes the geometry. A full four-node mesh needs six links, but four Sparks provide a natural four-link loop:

spark-329f → spark-355d → spark-c049 → spark-3d38 → spark-329f

Opposite ranks have no cable between them. That is the part that makes a switchless four-node deployment look suspicious at first: ordinary NCCL setup likes to construct more than one collective pattern, and RoCE is not normal routed IP traffic. Linux ip_forward does not magically turn a middle Spark into a RoCE router.

The useful observation is that tensor-parallel all-reduce does not require a full mesh if the collective is deliberately constrained to the physical ring. Each rank only needs to send to its next neighbor and receive from its previous one. The two-hop case is handled as a relay through the intermediate rank.

The physical and logical topology

Each cable carries two subnets, one for each PCIe root behind the ConnectX‑7. The four physical links are:

Cable Cages PCIe root A PCIe root B
329f ↔ 355d p1 / p1 192.168.100.0/24 192.168.102.0/24
355d ↔ c049 p0 / p0 192.168.104.0/24 192.168.105.0/24
c049 ↔ 3d38 p1 / p0 192.168.106.0/24 192.168.107.0/24
3d38 ↔ 329f p1 / p0 192.168.101.0/24 192.168.103.0/24

The rank order follows the cable order. That sounds obvious, but it is load-bearing: the ring implementation defines rank r - 1 as “previous” and rank r + 1 as “next.” A rank order that does not match the actual cables gives the collective a neighbor it cannot reach.

The live fabric view during inference. All four GB10 GPUs are at 96% utilization while traffic moves in both directions on every physical edge.
The live fabric view during inference. All four GB10 GPUs are at 96% utilization while traffic moves in both directions on every physical edge.

There are two different network stories in this setup:

  • RoCE data-plane traffic stays strictly neighbor-to-neighbor. It is never routed by Linux.
  • TCP control traffic—NCCL bootstrap, Gloo, and ZeroMQ decode-step coordination—can cross to the opposite rank through static routes and an explicit forwarding rule.

That separation matters. Treating RoCE like normal IP was one of the tempting dead ends.

How the collective works

I used the TP=RING4 path from kindlingai/glm-5.3-flash-gx10, pinned at commit 45b438b. The launcher tells the image which interfaces face the previous and next ranks. Its fabric setup then builds three pieces:

1. NCCL is forced to the Ring algorithm with a hand-written channel graph. Each channel receives on the previous port and sends on the next port. 2. arx handles the small all-reduces on direct neighbor queue pairs. At world size four, the proxy forwards the partial from the previous rank to the next rank. 3. arxbig handles larger reduce-scatter and all-gather traffic, splitting the opposite-rank path across both directions of the ring.

The essential environment is:

NCCL_ALGO=Ring
NCCL_IB_MERGE_NICS=0
NCCL_CROSS_NIC=1
NCCL_IB_SUBNET_AWARE_ROUTING=1
NCCL_IB_SUBNET_PREFIX_LEN=24
VLLM_ARX_RING=1

The important discovery on my real cables was NCCL_IB_SUBNET_AWARE_ROUTING=1. The upstream graph by itself failed: stock NCCL selected one device per channel for both directions, so a sender sometimes tried to reach its next neighbor through the port physically connected to its previous neighbor. Every rank eventually reported ibv_modify_qp ... Connection timed out.

With subnet-aware routing and a /24 prefix, the sender chooses the device that shares the receiver’s cable subnet. The same NCCL 2.30.7 build then passed from 1 MiB through 256 MiB, reaching 190 Gb/s bus bandwidth at 256 MiB. No patched NCCL was required, and both PCIe roots carried traffic.

Before starting vLLM I tested the transports in isolation:

Test Result
arx correctness 480 checks per rank, 0 failures
arx all-reduce 15.5 µs at 8 KB, 27.0 µs at 64 KB, 97.4 µs at 512 KB
arxbig all-gather 177–180 Gb/s at 8–32 MB
arxbig reduce-scatter 112–115 Gb/s at 8–32 MB
NCCL ring correct at 1–256 MiB, 190 Gb/s bus bandwidth at 256 MiB

First boot

The first boot took 717 seconds. That included 342 seconds to load the checkpoint, 99.5 seconds to write 45.3 GiB of tensor-parallel snapshots on every node, and 69 seconds for CUDA graph capture. Later starts can reuse the snapshots; the first boot is the expensive one.

The model uses the stock GLM‑5.3‑Flash NVFP4 checkpoint and stock drafter. Unlike TP=3, TP=4 needs no padded checkpoint: the attention heads, experts, and vocabulary divide cleanly by four.

With a 26 GiB KV-cache pin, vLLM exposed 4,241,096 KV tokens—enough for four simultaneous requests at roughly one million tokens each. That is 2.3× the 1.85M-token pool I had on TP=3.

TP=4 ring versus TP=3 mesh

I ran the same detached test battery against both cards: count-to-100, corruption probes, tool use, multi-turn prefix reuse, video input, long-context needle searches, and two repetitions of the decode benchmark twelve minutes apart.

The complete comparison from the test battery. Both TP=4 runs were taken from the same live card, twelve minutes apart.
The complete comparison from the test battery. Both TP=4 runs were taken from the same live card, twelve minutes apart.
Workload TP=3 mesh TP=4 ring Change
Count decode 128.9 tok/s 177.9–180.2 tok/s +39%
Code decode 94.5 tok/s 126.6–138.8 tok/s +34–47%
English prose decode 49.1 tok/s 68.6–70.3 tok/s +40–43%
Spanish prose decode 55.1 tok/s 71.7–72.3 tok/s +30%
Two-stream prose aggregate 73.2 tok/s 99.6–103.7 tok/s +36–42%
Cold prefill, 21.9K prompt 2,938 tok/s 4,499 tok/s +53%
Turn 1 / turn 2 TTFT 6.24 s / 1.72 s 4.13 s / 0.80 s faster
Needle near 1M 421.5 s 316.4 s faster
KV pool 1.85M tokens 4.24M tokens 2.3×

The ring is not “almost as good as a switch” in some abstract sense. On this model and this hardware it is fast enough that the fourth rank more than pays for the relay cost. The small-collective latency rises—8 KB arx all-reduce went from 12.2 µs on the three-node mesh to 15.5 µs on the four-node ring—but each rank holds 25% of the model rather than 33%, and the rest of the step gets faster.

A live code-generation window: 121.4 tok/s, 80.1% draft acceptance, and 6.03 accepted tokens per engine step.
A live code-generation window: 121.4 tok/s, 80.1% draft acceptance, and 6.03 accepted tokens per engine step.

The one-shot test that mattered more to me

Synthetic prompts are useful because they are repeatable, but I also wanted a task that looked like actual work. I asked the model:

Generate me in a single HTML file a landing page for an AI technology blog/personal branding site, name Carlos Alvarado, professional design, JS, CSS.

It produced a complete 54,214-token HTML document in one response. The run was:

Phase Tokens Time Rate
Prompt 44 0.822 s 53.6 tok/s
Generation 54,214 674.46 s 80.4 tok/s
Total 54,258 675.28 s —

That is eleven minutes and fifteen seconds for a large, self-contained page with responsive CSS and JavaScript—not a short code completion. The output is embedded below exactly as generated.

Interactive one-shot output — scroll inside the frame; links open separately.

Quality gates

Speed is only interesting if the fourth rank does not quietly change the model. The full battery passed:

  • count-to-100 exact output, 3/3, with 0.88–0.95 draft acceptance;
  • Hangul corruption probe, 0 replacement characters over three passes;
  • tool-call probe;
  • video-input probe using frames across time;
  • upstream smoke test, 8/8;
  • needle retrieval near 300K, 620K, and 1M tokens;
  • no error line on any rank while serving.

The one warning worth recording was a single arx ... STALL message on every rank during CUDA graph capture. Boot continued immediately, and no stall appeared under serving traffic.

At the one-million-token needle, the head had only 3.2 GiB free. Workers retained 6.5–9.7 GiB. The head is therefore the tight point of the card, and I keep speech workloads on a worker.

What remains open

This is a working deployment, not the end of the investigation.

  • The first validation run covered roughly 30 minutes of mixed traffic. A longer soak still matters because one upstream report saw a rank die after 1.5 hours on a later commit and a different RecoverSSM configuration.
  • Snapshot-restore boot time still needs a clean measurement.
  • TP=3 cannot run unchanged on this ring. Any three of the four nodes form a line, not a closed collective. Switching back to the triangle requires changing the cables and applying the mesh layout.

Those limits are operational, not hypothetical excuses. The card is live, has passed the quality battery, and has already generated long real outputs correctly.

The practical conclusion

For weeks I treated the absence of published four-Spark ring results as evidence that a switch was probably required. It was not. The working recipe was already latent in the upstream ring code; it needed the physical rank order, per-cable subnets, the right NCCL device selection on real cables, and separate treatment of RoCE and TCP.

The result is now my preferred GLM‑5.3‑Flash configuration:

  • four DGX Sparks;
  • direct ConnectX‑7 cables in a ring;
  • no switch;
  • stock NCCL 2.30.7;
  • GLM‑5.3‑Flash NVFP4 with DFlash2;
  • 68.6–70.3 tok/s prose, 126.6–138.8 tok/s code;
  • 4.24M KV tokens;
  • verified text, tools, long context, and video.

The biggest lesson is simple: a topology that is not widely documented is not the same thing as a topology that cannot work. Measure the actual transport, force the collective to match the cables, and test correctness independently from serving. Four Sparks on a ring are not a consolation prize. For this model, they are the fastest and most useful configuration I have run.


Code and reproducibility: calvarado2004/sparks/glm53-kindling · bench/glm-battery.sh · upstream report: kindlingai/glm-5.3-flash-gx10#88