bifrost
View on GitHubFastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.
Bifrost is a high-performance Go AI gateway exposing one OpenAI-compatible API over 23+ providers, with automatic failover, key load balancing, semantic caching, MCP tool support, guardrails and budget governance at low overhead.
Use Cases
Single OpenAI-compatible API across 23+ LLM providersAutomatic failover and retries between providers/modelsLoad balancing requests across multiple API keysSemantic caching to cut LLM cost and latencyMCP client/server gateway for model tool callingBudget, virtual key and rate-limit governanceLLM observability, tracing and Prometheus metricsDrop-in replacement for OpenAI/Anthropic/GenAI base URLsMultimodal and streaming request handlingSelf-hosted Go SDK embedded gateway deployment
Built With
- Language
- Go
- Frameworks
- LangChain · LiteLLM · OpenAI SDK · Anthropic SDK · AWS Bedrock · Google Vertex AI · Google GenAI SDK · Ollama · Groq · Mistral · Cohere · Cerebras · Prometheus · Docker · Helm · Kubernetes
Tags
ai-gateway · llm-gateway · model-router · load-balancing · guardrails · semantic-cache · mcp · observability · llmops · failover · multi-provider · openai-compatible · go · governance · budget-management · streaming