ome
★ 513OME is a Kubernetes operator for LLM serving: model-as-CRDs, GPU bin-packing scheduling, and automatic runtime selection across SGLang, vLLM, TensorRT-LLM and Triton. Adds prefill-decode disaggregation, multi-node inference, LoRA serving, autoscaling and benchmarking.
AI Frameworks | Go · kubernetes · llm-serving
View Project →