vllm-mlx
★ 1.6KvLLM-style inference server for Apple Silicon built on MLX, exposing OpenAI /v1/* and Anthropic /v1/messages from one process. Adds continuous batching, paged/prefix KV cache, structured output, MCP tool calling, and multimodal text, vision, audio, embeddings and rerank support.
AI Frameworks | Python · llm-inference · mlx
View Project →