代码库

A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that share a prompt reuse each other's KV cache instead of each prefilling it. Rust, wrapping llama.cpp.
Rust
cudaggufinference-enginellama-cppllmlocal-llmmlxollamaollama-apiopenai-apirocmrustvulkan