logo
GLM-5.2 on a 4× GB10 cluster: ~22 tok/s decode, 256K ctx, Recipe - DGX Spark / GB10 - NVIDIA Developer Forums

GLM-5.2 on a 4× GB10 cluster: ~22 tok/s decode, 256K ctx, Recipe - DGX Spark / GB10 - NVIDIA Developer Forums

Got GLM-5.2 (GlmMoeDsa / DeepSeek-Sparse-Attention arch) serving on 4× GB10 over a 100G MikroTik CRS504 switch. It was quite the hassle to get DSA + MTP running, but then again Claude did 90% of the work : ) Model cyank…

Related Recipes