Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
SMG is a Rust LLM gateway that unifies OpenAI/Anthropic/Gemini and self-hosted engines (vLLM, SGLang, TensorRT-LLM, TokenSpeed, MLX) behind one API. It adds KV-cache-aware routing, gRPC pipelines, MCP tooling, multi-tenancy and observability for large-scale inference deployments.
Use Cases
Built With
- Language
- Rust
- Frameworks
- vLLM · SGLang · TensorRT-LLM · TokenSpeed · MLX · Ollama · Tokio · Axum · Tonic · Kubernetes · Helm · Docker · OpenTelemetry · Prometheus · PostgreSQL · Redis
Tags
llm-gateway · rust · inference-routing · load-balancing · kv-cache-aware · openai-compatible · anthropic-api · grpc · vllm · sglang · tensorrt-llm · mcp · wasm-plugins · multi-tenant · observability · embeddings