Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
OME is a Kubernetes operator for LLM serving: model-as-CRDs, GPU bin-packing scheduling, and automatic runtime selection across SGLang, vLLM, TensorRT-LLM and Triton. Adds prefill-decode disaggregation, multi-node inference, LoRA serving, autoscaling and benchmarking.
Use Cases
Deploy and serve LLMs on Kubernetes as custom resourcesAutomatic runtime selection between SGLang, vLLM, TensorRT-LLM and TritonGPU bin-packing and accelerator-aware schedulingPrefill-decode disaggregated and multi-node inferenceLoRA adapter and fine-tuned weight servingModel lifecycle, storage and encryption managementCustom-metrics autoscaling for inference servicesBenchmarking model performance under configurable load
Built With
- Language
- Go
- Frameworks
- SGLang · vLLM · TensorRT-LLM · Triton · Kubernetes · Kueue · LeaderWorkerSet · KEDA · Gateway API · Istio · Prometheus · Volcano · controller-runtime · Knative · Helm
Tags
kubernetes · llm-serving · model-serving · gpu-scheduling · inference · kubernetes-operator · multi-node · autoscaling · prefill-decode-disaggregation · lora · model-lifecycle · helm · go · benchmarking · mcp-gateway · gpu