THE FIRST DLSS 5 MANAGER
-
Updated
Sep 28, 2026 - F#
THE FIRST DLSS 5 MANAGER
Cross-platform installer for Triton and SageAttention on ComfyUI. Simplifies GPU-accelerated inference setup for Windows users with automated dependency management and RTX 5090 support.
Tiny local text-to-24x24 pixel art model, trained on roughly 200K samples in 30 minutes on an RTX 5090.
Fastest MoE/LLM inference runtime for consumer and edge Blackwell GPUs. SN74 on Gittensor.
Native Windows port of NInfer engine for RTX 5090. Features Qwen3.8-27B with QUASAR and NInfer models, MTP/DFlash2 with vision and 262,144 context
LLM server for one RTX 5090. Runs Qwen3.8-Flash-Next (512-expert MoE) on a single 32 GB card at 75-80 tok/s. Decodes 17-191 % faster than llama.cpp on the same models. Built for agents: tool calls, long context, many streams at once.
Crazy Fast Qwen3.8–27b With a 420k Context on an RTX 5090
Durable local inference for Oh My Pi: NInfer + Qwen3.8 27B on one RTX 3090/4090/5090, with restart-resumable OpenAI Responses state.
RTX 5090 & RTX 5060 Docker container with PyTorch + TensorFlow. First fully-tested Blackwell GPU support for ML/AI. CUDA 12.8, Python 3.11, Ubuntu 24.04. Works with RTX 50-series (5090/5080/5070/5060) and RTX 40-series.
异环(Neverness To Everness / Ananta)光线追踪一键部署面板��基于 OptiScaler winmm 方案,默认推荐 RTX 5090,并支持本机/RTX 4090/RTX 5080M 配置、备份、恢复和本地 WebUI。
DLSS 5 Neural Rendering for World of Warcraft (WoW) on NVIDIA RTX 40/50 — a safe overlay with no DLL injection, no hooks and no ReShade next to Wow.exe. Free and open source.
NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.
Research: vGPU unlock on consumer NVIDIA RTX 5090 (Blackwell/GB202). 19 binary patches, full CPU-side pipeline working, GSP firmware blocked by fused-off VF PRIV registers.
Pixal3D ComfyUI integration for Windows (RTX 30/40/50) — single image to textured PBR mesh in 3-5 min
Qwen3.8-27B on RTX 5090s — 262K ctx, 1.4M-token KV pool, ~220 t/s code decode. NVFP4 + vLLM + sm120 patches, reproducible.
Windows prebuilt of llama.cpp combining Multi-Token Prediction (MTP) + TurboQuant KV cache compression + native sm_120 (Blackwell consumer GPU, FP4 tensor cores). For RTX 5060 Ti / 5070 / 5080 / 5090.
Qwen3.8-27B TWIN-TURBO (NVFP4 GGUF, all-Q8_0 precision chain) tuned to its limits on RTX 5090 Laptop 24GB: 74-78 tok/s with built-in MTP + CUDA graphs, overthinking -93%, vision 3.9s, 160K stable context (cliff mapped). 10 rounds of measured tuning, rejected approaches documented, continuously updated.
Per-pin 12VHPWR monitoring for ASUS ROG Astral cards as standard Linux hwmon sensors, plus astral-guard
To associate your repository with the rtx-5090 topic, visit your repo's landing page and select "manage topics."