vllm
View on GitHubA high-throughput and memory-efficient inference and serving engine for LLMs
vLLM is a high-throughput, memory-efficient LLM inference and serving engine built on PagedAttention, continuous batching, and optimized CUDA/ROCm kernels. It exposes an OpenAI-compatible server with quantization, LoRA, speculative decoding, and disaggregated serving across 200+ model architectures.
Use Cases
high-throughput LLM inference and servingOpenAI-compatible API endpoint for self-hosted modelsoffline batch and throughput benchmarkingmulti-LoRA adapter servingembedding and retrieval model serving for RAGstructured/JSON output generation via xgrammartool calling and reasoning parser pipelinesspeculative decoding for latency reductiondisaggregated prefill/decode deploymentquantized model deployment (FP8, INT4, GPTQ/AWQ, GGUF)multimodal vision-language model servingreward and classification model scoringdistributed multi-GPU/multi-node inference
Built With
- Language
- Python
- Frameworks
- PyTorch · CUDA · HIP/ROCm · Triton · CUTLASS · FlashAttention · FlashInfer · xgrammar · guidance · Hugging Face Transformers · Hugging Face Hub · torch.compile · Ray · gRPC
Tags
inference-engine · llm-serving · paged-attention · continuous-batching · quantization · speculative-decoding · moe · openai-compatible-api · kv-cache · distributed-inference · lora · structured-output · cuda · rocm · tpu · throughput