Github

代码库

Qwen3.8-27B on one RTX 4090: full native 262K context via E8 4-bit KV, up to 149 tok/s code decode with MTP3, sm_89-retuned attention prefill, vision, llama.cpp-compatible /metrics + /slots
C++
cudainference-enginellmqwenrtx-4090