vllm-mlx
View on GitHubHigh-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
vLLM-style inference server for Apple Silicon built on MLX, exposing OpenAI /v1/* and Anthropic /v1/messages from one process. Adds continuous batching, paged/prefix KV cache, structured output, MCP tool calling, and multimodal text, vision, audio, embeddings and rerank support.
Use Cases
Run a local OpenAI-compatible LLM API on a MacServe a local Anthropic /v1/messages endpoint for Claude CodeHigh-throughput multi-user LLM serving with continuous batchingMultimodal image/video/audio chat inferenceLocal text-to-speech and speech-to-textEmbeddings and reranking for RAG pipelinesStructured JSON output via response_formatReasoning model output extraction (Qwen3, DeepSeek-R1)Benchmarking inference throughput and latencyLong-context agents with SSD-tiered KV cache
Built With
- Language
- Python
- Frameworks
- MLX · mlx-lm · mlx-vlm · mlx-audio · mlx-embeddings · FastAPI · Uvicorn · Starlette · Hugging Face Transformers · PyTorch · Gradio · MCP SDK · lm-format-enforcer · Prometheus
Tags
llm-inference · mlx · apple-silicon · openai-compatible · anthropic-api · continuous-batching · kv-cache · multimodal · mcp · tool-calling · text-to-speech · speech-to-text · embeddings · reranking · claude-code · local-llm