Vibe Coding Discover

Use Cases

Quantized Inference (fp4/fp8/int4/awq/gptq)

Published projects tagged with this use case.

1 project

sglang

★ 36K

SGLang is a high-performance serving framework for LLMs and multimodal models, offering RadixAttention prefix caching, prefill-decode disaggregation, speculative decoding, and quantization over an OpenAI-compatible API across NVIDIA, AMD, TPU, and CPU hardware.

AI Frameworks | Python · llm-serving · inference-engine

View Project →