代码库

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Python
apple-siliconinference-serverllmmacosmlxopenai-api