Skip to content
View bemlerlabs's full-sized avatar

Block or report bemlerlabs

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. qwen3.8-flash-next-dgx-spark-sglang qwen3.8-flash-next-dgx-spark-sglang Public

    Production recipe for serving Qwen3.8-Flash-Next (180B Hybrid MoE) on a single 128GB NVIDIA DGX Spark with SGLang, NVMe PLE offloading & MTP speculative decoding.

    Jinja 9

  2. glm-5.3-flash-dgx-spark-dflash2 glm-5.3-flash-dgx-spark-dflash2 Public

    A practical recipe for serving GLM-5.3-Flash (320B MoE) on a single 128GB NVIDIA DGX Spark with DFlash2 speculative decoding.

    Shell 5