NInfer for Windows and around 16GB VRAM: RTX 5070 Ti / 5080 / 5090, Qwen3.8-27B GSQ-RCO Q3, CUDA 13 Native engine, tray manager, model conversion and measured setup guides.
-
Updated
Oct 1, 2026 - C++
NInfer for Windows and around 16GB VRAM: RTX 5070 Ti / 5080 / 5090, Qwen3.8-27B GSQ-RCO Q3, CUDA 13 Native engine, tray manager, model conversion and measured setup guides.
High-performance single-GPU inference for selected model checkpoints and GPUs.
NInfer fork with end-user improvements such as model router and jev-alike decisions endpoint. Experimental repo, changes are subject to be wiped without notice based on my own usage observations. Packaged with nix.
Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.
Tesla V100 32GB (sm_70) running Qwen3.8-27B: sm70 decode kernel port plus KV context-cache tuning, measured on a real 53-request agent session. Decode 42.3-89.4 tok/s, TTFT 0.54 s on a cache hit, 200k-token prompts, zero failed requests, raw engine logs included. Published by an AI on the machine owner's behalf. 中文版:README.zh-CN.md
Port NInfer, a single-GPU CUDA inference engine, to the NVIDIA L20 (Ada sm_89, 92 SMs, 48 GB): patch set, build tooling, and measured results
笔记本 RTX 5070 Ti Laptop 12G(12,227 MiB / 140 W TGP)· Bonsai-2-27B 三元量化 · MTP vs DFlash2 同上下文 A/B���Laptop GPU only — NOT the desktop 5070 Ti (16 GB / ~300 W); numbers are not comparable across the two.
Qwen3.8-27B on RTX 4090 D (48GB): production deployment of NInfer with MTP7 + E8 KV + NVMe disk cache, 195 tok/s decode, crash forensics for WDDM desktop GPUs
Rust CLI to run local LLMs through rootless Podman Compose
Measuring proxy + throughput dashboard for local LLM engines (llama.cpp, NInfer): live tok/s, cache-hit rate, TTFT, and history charts. Single-file panel, stdlib-only proxy, MIT.
Unofficial Windows distribution of iamwavecut/ninfer-all with target-specific builds and an integration contract for NInferEZ.
To associate your repository with the ninfer topic, visit your repo's landing page and select "manage topics."