Twitter/X

GLM 5.2 NVFP4 REAP 469B model is reported running on a 3x DGX Spark cluster.

Brief

GLM 5.2 NVFP4 REAP 469B is running on a 3× DGX Spark cluster, delivering a 256,000-token context window with about 4.4 tokens/second throughput. The implementation references prior work by @0xSero, and the repository for the GLM-spark setup is published at github.com/bird/GLM-spark.

Why it matters

GLM 5.2 NVFP4 REAP 469B model is reported running on a 3x DGX Spark cluster.

Key details

  • The setup supports a 256K token context window and achieves approximately 4.4 tokens/second throughput.
  • The work is credited as based on contributions from @0xSero and linked code is available at github.com/bird/GLM-spark.
Source evidence

GLM 5.2 NVFP4 REAP 469B running on 3X DGX Spark. 256K context @ ~4.4 tok/s. based on work done by @0xSero

github.com/bird/GLM-spark