Vibe Coding Discover

Use Cases

Custom-metrics Autoscaling For Inference Services

Published projects tagged with this use case.

1 project

ome

★ 513

OME is a Kubernetes operator for LLM serving: model-as-CRDs, GPU bin-packing scheduling, and automatic runtime selection across SGLang, vLLM, TensorRT-LLM and Triton. Adds prefill-decode disaggregation, multi-node inference, LoRA serving, autoscaling and benchmarking.

AI Frameworks | Go · kubernetes · llm-serving

View Project →