Hermes-router
View on GitHubOpenAI/Anthropic-compatible AI router that keeps apps online with provider failover, key rotation, caching, analytics, and local-model fallback.
A self-hosted Python gateway for OpenAI- and Anthropic-compatible requests. It routes across cloud and local model providers with failover, key rotation, caching, budgets, and a monitoring dashboard.
Use Cases
Route OpenAI- or Anthropic-compatible app requests across multiple providersFail over automatically when a provider is unavailable or rate-limitedUse local Ollama or LM Studio models with cloud-provider fallbackChoose lower-cost models based on request difficulty and model capabilityMonitor usage, latency, provider health, and spending
Built With
- Language
- Python
- Frameworks
- Flask · Waitress · LangChain
Tags
LLM router · API proxy · provider failover · key rotation · model routing · caching · usage analytics · rate limiting · local models · OpenAI-compatible · Anthropic-compatible · embeddings · tool calling · Prometheus