sglang
★ 36KSGLang is a high-performance serving framework for LLMs and multimodal models, offering RadixAttention prefix caching, prefill-decode disaggregation, speculative decoding, and quantization over an OpenAI-compatible API across NVIDIA, AMD, TPU, and CPU hardware.
AI Frameworks | Python · llm-serving · inference-engine
View Project →