Vibe Coding Discover

AI Frameworks

SGLang is a high-performance serving framework for large language models and multimodal models.

★ 36K9,097 forksPythonApache-2.0sgl-project

SGLang is a high-performance serving framework for LLMs and multimodal models, offering RadixAttention prefix caching, prefill-decode disaggregation, speculative decoding, and quantization over an OpenAI-compatible API across NVIDIA, AMD, TPU, and CPU hardware.

Use Cases

High-throughput LLM serving behind an OpenAI-compatible APILow-latency chat/completion inference at scaleServing vision-language and multimodal modelsImage and video diffusion generation workflowsRL rollout backend for post-trainingStructured JSON / constrained output decodingMulti-LoRA adapter batching on shared GPUsMulti-node distributed inference with tensor/expert parallelismPrefix-cache reuse for long shared promptsQuantized inference (FP4/FP8/INT4/AWQ/GPTQ)Embedding and reward model serving

Built With

Language
Python
Frameworks
PyTorch · CUDA · ROCm · Triton · FlashInfer · sgl-kernel · Rust · Ray · JAX · Hugging Face Transformers · vLLM · OpenAI API

Tags

llm-serving · inference-engine · radixattention · prefix-caching · speculative-decoding · quantization · moe · multimodal · vlm · diffusion · distributed-inference · openai-compatible · reinforcement-learning · continuous-batching · paged-attention · high-throughput

sglang — Vibe Coding Discover