MTPLX
★ 2.5KMTPLX is an Apple Silicon LLM inference engine and Mac app that uses Qwen's native multi-token prediction heads for exact speculative decoding (1.6x-2.24x faster decode at any temperature). It serves local models over OpenAI- and Anthropic-compatible APIs for coding agents and chat.
AI Frameworks | Python · apple-silicon · mlx
View Project →