Skip to content
View 0xBakeer's full-sized avatar
💭
https://x.com/0xbakeer
💭
https://x.com/0xbakeer

Block or report 0xBakeer

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. TandemLLM TandemLLM Public

    Inference engine for Qwen3.8-27B on one DGX Spark: speculative draft trees sized by StairCut, NVFP4 kernels, exact recurrent-state caches

    Python 21

  2. inference-atlas inference-atlas Public

    A community-owned map of LLM inference engine configurations — benchmarks, evals and coverage gaps, hosted on GitHub Pages. The repo is the database.

    Python 16 6

  3. arbiter arbiter Public

    Serve typed-decision (System 1) models — Laya or your own — on NVIDIA GPUs or Apple Silicon, with a Jev-compatible API and coding-agent integrations

    HTML 33 3

  4. qwen38-flash-next-spark qwen38-flash-next-spark Public

    Run Qwen3.8-Flash-Next (180B) on a single DGX Spark by keeping its 51B n-gram embedding table on NVMe

    Shell 125 9

  5. deepseek-v41-flash-spark deepseek-v41-flash-spark Public

    DeepSeek-V4.1-Flash on a single DGX Spark (GB10): resident hot experts + NVMe streaming, DSpark, OpenAI API. Work in progress.

    Python 120 13