Skip to content
#

ninfer

Here are 16 public repositories matching this topic...

Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.

  • Updated Sep 29, 2026
  • Cuda
ninfer-v100-sm70-decode

Tesla V100 32GB (sm_70) running Qwen3.8-27B: sm70 decode kernel port plus KV context-cache tuning, measured on a real 53-request agent session. Decode 42.3-89.4 tok/s, TTFT 0.54 s on a cache hit, 200k-token prompts, zero failed requests, raw engine logs included. Published by an AI on the machine owner's behalf. 中文版:README.zh-CN.md

  • Updated Oct 1, 2026
  • Shell

Add this topic to your repo

To associate your repository with the ninfer topic, visit your repo's landing page and select "manage topics."

Learn more