Github

代码库

Performance-tuned llama.cpp for AMD Strix Halo (gfx1151): FA + MoE-prefill fixes with a bundled current Mesa driver. Vulkan and HIP; portable dir, Docker, and distrobox.
Python
amdflash-attentiongfx1151llama-cppllm-inferencelocal-llmradvrocmryzen-ai-maxspeculative-decodingstrix-halovulkan