Vibe Coding Discover

AI Tools

Switchyard

View on GitHub

Switchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility - enabling flexible model selection, benchmarking, and cost/performance optimization.

★ 3.2K309 forksRustApache-2.0NVIDIA-NeMo

NVIDIA Switchyard is a Rust+Python LLM routing layer that picks the cheapest capable model per call while staying OpenAI/Anthropic API compatible. It works as a NeMo Relay plugin, a LiteLLM routing plugin, an embeddable library (switchyard-libsy), or a standalone proxy that Claude Code and Codex CLI can point at.

Use Cases

Route each LLM call to the cheapest model that can still complete the taskCut inference cost for coding agents without changing agent codeRun an OpenAI/Anthropic-compatible proxy in front of Claude Code or Codex CLIBenchmark model accuracy versus total token cost on Terminal-BenchEmbed dynamic model routing into an existing LLM gateway or harnessEscalate hard requests from an efficient model to a capable oneUse an LLM classifier to pick a model per request by task capabilityExport Prometheus metrics for requests, latency, tokens, and routing overheadLoad routing rules into an existing NeMo Relay deployment via routes.tomlAdd a routing plugin to a LiteLLM Router or proxy

Built With

Language
Rust
Frameworks
LiteLLM · NeMo Relay · OpenAI API · Anthropic Messages API · OpenAI Responses API · PyO3 · Maturin · Prometheus · MkDocs

Tags

llm-routing · model-routing · llm-gateway · openai-compatible · anthropic-compatible · cost-optimization · proxy-server · litellm · provider-abstraction · routing-algorithms · rust · python · benchmarking · prometheus-metrics · claude-code · codex-cli