Github

代码库

A high-performance and light-weight router for vLLM large scale deployment
Rust
Agent skills for vLLM
Shell
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Python
vLLM Quantization plugin for GGUF
Python
Community maintained hardware plugin for vLLM on Apple Silicon
Python
apple-siliconllmmacosmetalmlxvllm
Cost-efficient and pluggable Infrastructure components for GenAI inference
Go
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Python
compressionquantization
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
Python
A programmable Mixture-of-Models router for heterogeneous LLM inference
Go
ai-gatewayguardrailsinferencekubernetesllmllmroutermixture-of-modelspytorchsemantic-routertransformervllm
Common recipes to run vLLM
JavaScript
A framework for efficient model inference with omni-modality models
Python
audio-generationdiffusionimage-generationinferencemodel-servingmultimodalpytorchtransformervideo-generationworld-model
A high-throughput and memory-efficient inference and serving engine for LLMs
Python
amdblackwellcudadeepseekdeepseek-v3gptgpt-ossinferencekimillamallmllm-servingmodel-servingmoeopenaipytorchqwenqwen3tputransformer
Community maintained hardware plugin for vLLM on Huawei Ascend
C++
ascendinferencellmllm-servingllmopsmlopsmodel-servingtransformervllm