Switchyard
View on GitHubSwitchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility - enabling flexible model selection, benchmarking, and cost/performance optimization.
NVIDIA Switchyard is a Rust+Python LLM routing layer that picks the cheapest capable model per call while staying OpenAI/Anthropic API compatible. It works as a NeMo Relay plugin, a LiteLLM routing plugin, an embeddable library (switchyard-libsy), or a standalone proxy that Claude Code and Codex CLI can point at.
Use Cases
Built With
- Language
- Rust
- Frameworks
- LiteLLM · NeMo Relay · OpenAI API · Anthropic Messages API · OpenAI Responses API · PyO3 · Maturin · Prometheus · MkDocs
Tags
llm-routing · model-routing · llm-gateway · openai-compatible · anthropic-compatible · cost-optimization · proxy-server · litellm · provider-abstraction · routing-algorithms · rust · python · benchmarking · prometheus-metrics · claude-code · codex-cli