Vibe Coding Discover

Use Cases

Low-latency Chat/completion Inference At Scale

Published projects tagged with this use case.

1 project

sglang

★ 36K

SGLang is a high-performance serving framework for LLMs and multimodal models, offering RadixAttention prefix caching, prefill-decode disaggregation, speculative decoding, and quantization over an OpenAI-compatible API across NVIDIA, AMD, TPU, and CPU hardware.

AI Frameworks | Python · llm-serving · inference-engine

View Project →