Skip to content
#

rtx-5090

Here are 115 public repositories matching this topic...

LLM server for one RTX 5090. Runs Qwen3.8-Flash-Next (512-expert MoE) on a single 32 GB card at 75-80 tok/s. Decodes 17-191 % faster than llama.cpp on the same models. Built for agents: tool calls, long context, many streams at once.

  • Updated Oct 2, 2026
  • C++

异环(Neverness To Everness / Ananta)光线追踪一键部署面板��基于 OptiScaler winmm 方案,默认推荐 RTX 5090,并支持本机/RTX 4090/RTX 5080M 配置、备份、恢复和本地 WebUI。

  • Updated May 18, 2026
  • Python

Qwen3.8-27B TWIN-TURBO (NVFP4 GGUF, all-Q8_0 precision chain) tuned to its limits on RTX 5090 Laptop 24GB: 74-78 tok/s with built-in MTP + CUDA graphs, overthinking -93%, vision 3.9s, 160K stable context (cliff mapped). 10 rounds of measured tuning, rejected approaches documented, continuously updated.

  • Updated Sep 28, 2026
  • Python

Add this topic to your repo

To associate your repository with the rtx-5090 topic, visit your repo's landing page and select "manage topics."

Learn more