magnitude
View on GitHubOpen source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or just a CPU.
A local inference engine that compiles and tunes kernels for the host hardware to run open models. Includes a desktop app and CLI, and connects to agents such as Pi, OpenCode, Hermes, and Codex.
Use Cases
Run open models locally for AI agentsServe models through an OpenAI-compatible APIOptimize inference for Apple Silicon, NVIDIA, AMD, or CPU hardwareConnect existing agents to local inference
Built With
- Language
- Rust
Tags
LLM inference · local AI · open models · kernel compilation · hardware tuning · GPU acceleration · CPU inference · OpenAI-compatible API · agent inference