# Glm 5.3 Flash 4x Dgx Spark Switchless

> GLM-5.3-Flash (NVFP4) at TP4 across 4x DGX Spark via a switchless RoCE ring + DFlash2 — reproducible recipe - alexellis/glm-5.3-flash-4x-dgx-spark-switchless

- **Website:** https://github.com/alexellis/glm-5.3-flash-4x-dgx-spark-switchless
- **Pricing:** unknown
- **Categories:** Developer Tools, AI Agents
- **Tags:** developer-tools, ai-agents, sales-marketing
- **Platforms:** Web
- **Last verified:** 2026-09-10
- **Canonical page:** https://linkrena.com/tools/glm-5-3-flash-4x-dgx-spark-switchless

## About

GLM-5.3-Flash NVFP4 — 4× DGX Spark, switchless-ring TP4 + DFlash2

This repository is the recipe and the contract : every address, interface, and hostname is a placeholder you swap for your own — nothing here depends on a private gateway, router, or network.

Not tied to these weights, either. The ring fabric, the patched NCCL, and the rank-launch pattern know nothing about the model — the same four-node fabric has also served GLM-5.2 (in more than one quant format: EXL3, and QuantTrio) and DeepSeek-V4-Flash at TP4. To adapt it, swap the weights and the model-specific serve arguments (parsers, drafter, MoE backend, and KV sizing); everything else in the recipe carries over.

And no — this is not Sparkring under a different name. There is no Sparkring/SIRCL anywhere in this stack. The collectives form on a patched NCCL 2.30.7 — a skip-tree-connect change LD_PRELOAD -ed into every container, because stock NCCL's tree-connect step wedges on a switch-free point-to-point fabric — plus a pinned NCCL runtime profile (ring algorithm, fixed channels, subnet-aware dual-rail routing) worked out for this topology. That combination was validated on this ring with Sparkring absent entirely.

## Related tools

- [mabl](https://linkrena.com/tools/mabl): The #1 agentic testing platform empowering software teams to accelerate releases, ensure software quality, and deliver exceptional user experiences.
- [Experiential](https://linkrena.com/tools/experiential): An open source model gateway that provides one control plane across closed, open-source, local, and custom models.
- [Talos](https://linkrena.com/tools/talos): Hand it a shell. Every tool call passes a gate that grants authority for exact arguments, once, for 30 seconds — and logs the verdict.
- [Metis](https://linkrena.com/tools/metis): Metis is a coding agent that boosts AI/LLM coding performance by 50% - Wholiver/metis
- [Slotstream](https://linkrena.com/tools/slotstream): Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD.
- [Senator Bernie Sanders](https://linkrena.com/tools/senator-bernie-sanders): WASHINGTON, Sept. 3 — Sen.
