Rapid-MLX
★ 3.9KLocal LLM inference server and macOS app for Apple Silicon, built on MLX. Provides OpenAI- and Anthropic-compatible APIs, continuous batching, and tool-calling support for coding agents.
AI Frameworks | Python · LLM inference · Apple Silicon
View Project →