Vibe Coding Discover

AI Frameworks

Run LLMs with MLX

★ 7.2K1,084 forksPythonMITml-explore

MLX LM is a Python package for running and fine-tuning LLMs on Apple silicon via MLX. It offers CLI and Python APIs for generation, chat, LoRA/full fine-tuning, quantization, GGUF conversion, prompt caching, and an OpenAI-compatible server.

Use Cases

Run LLMs locally on Apple siliconFine-tune models with LoRA or full fine-tuningQuantize and upload models to Hugging Face HubServe an OpenAI-compatible local inference serverInteractive chat REPL with a local modelCache long prompts for reuse across queriesDistributed inference and fine-tuning across MacsConvert HF models to MLX formatEvaluate perplexity and benchmark throughputStream token generation via a Python API

Built With

Language
Python
Frameworks
MLX · Hugging Face Hub · Transformers · safetensors · LoRA · GGUF · AWQ · GPTQ

Tags

llm-inference · apple-silicon · mlx · quantization · fine-tuning · lora · gguf · text-generation · prompt-caching · local-llm · openai-compatible · cli · chat-repl · model-conversion · distributed-inference · benchmarking