FreeToken
View on GitHubFreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
FreeToken is a local MoE inference and serving engine that runs large open-weight models across GPU, CPU, and host memory. It provides OpenAI- and Anthropic-compatible APIs, plus a desktop app and CLI.
Use Cases
Run large MoE models locally on consumer hardwareServe local models through OpenAI-compatible APIsServe local models through Anthropic-compatible APIsPower coding and tool-calling agents with local inference
Built With
- Language
- Python
- Frameworks
- PyTorch · FastAPI · Hugging Face Hub · GGUF
Tags
LLM inference · MoE serving · local inference · edge AI · GPU · CPU offloading · OpenAI-compatible API · Anthropic-compatible API · model serving