Vibe Coding Discover

AI Frameworks

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 92K22,526 forksPythonApache-2.0vllm-project

vLLM is a high-throughput, memory-efficient LLM inference and serving engine built on PagedAttention, continuous batching, and optimized CUDA/ROCm kernels. It exposes an OpenAI-compatible server with quantization, LoRA, speculative decoding, and disaggregated serving across 200+ model architectures.

Use Cases

high-throughput LLM inference and servingOpenAI-compatible API endpoint for self-hosted modelsoffline batch and throughput benchmarkingmulti-LoRA adapter servingembedding and retrieval model serving for RAGstructured/JSON output generation via xgrammartool calling and reasoning parser pipelinesspeculative decoding for latency reductiondisaggregated prefill/decode deploymentquantized model deployment (FP8, INT4, GPTQ/AWQ, GGUF)multimodal vision-language model servingreward and classification model scoringdistributed multi-GPU/multi-node inference

Built With

Language
Python
Frameworks
PyTorch · CUDA · HIP/ROCm · Triton · CUTLASS · FlashAttention · FlashInfer · xgrammar · guidance · Hugging Face Transformers · Hugging Face Hub · torch.compile · Ray · gRPC

Tags

inference-engine · llm-serving · paged-attention · continuous-batching · quantization · speculative-decoding · moe · openai-compatible-api · kv-cache · distributed-inference · lora · structured-output · cuda · rocm · tpu · throughput