Vibe Coding Discover

AI Frameworks

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.

★ 543173 forksRustApache-2.0smg-project

SMG is a Rust LLM gateway that unifies OpenAI/Anthropic/Gemini and self-hosted engines (vLLM, SGLang, TensorRT-LLM, TokenSpeed, MLX) behind one API. It adds KV-cache-aware routing, gRPC pipelines, MCP tooling, multi-tenancy and observability for large-scale inference deployments.

Use Cases

Route LLM traffic across self-hosted engines and cloud providers behind one endpointKV-cache-aware prefix routing to maximize GPU utilizationDrop-in OpenAI and Anthropic API compatibility for any backendPrefill/decode disaggregation and DP-aware routing for inference enginesMulti-tenant API-key auth with priority admission schedulingMCP tool discovery and execution with approval policiesPersistent chat-history storage with pluggable backendsExtend request/response handling with WASM middleware pluginsKubernetes-native worker discovery and HA mesh networkingObservability via 90+ Prometheus metrics and OpenTelemetry tracing

Built With

Language
Rust
Frameworks
vLLM · SGLang · TensorRT-LLM · TokenSpeed · MLX · Ollama · Tokio · Axum · Tonic · Kubernetes · Helm · Docker · OpenTelemetry · Prometheus · PostgreSQL · Redis

Tags

llm-gateway · rust · inference-routing · load-balancing · kv-cache-aware · openai-compatible · anthropic-api · grpc · vllm · sglang · tensorrt-llm · mcp · wasm-plugins · multi-tenant · observability · embeddings