Rapid-MLX
View on GitHubRapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server and Mac app for Apple Silicon, built on MLX, focused on reliable tool calling for coding agents. Release-gated with Claude Code, Codex CLI, Aider, Hermes and DeepSeek Harness.
Local LLM inference server and macOS app for Apple Silicon, built on MLX. Provides OpenAI- and Anthropic-compatible APIs, continuous batching, and tool-calling support for coding agents.
Use Cases
Serve local LLMs on Apple SiliconProvide a local inference backend for coding agentsRun concurrent LLM requests with continuous batchingUse OpenAI- or Anthropic-compatible APIs with local modelsTest tool calling across agent clients
Built With
- Language
- Python
- Frameworks
- MLX · mlx-lm
Tags
LLM inference · Apple Silicon · local AI · inference server · tool calling · continuous batching · OpenAI-compatible API · Anthropic-compatible API · speculative decoding · coding agents · macOS