colibri
View on GitHubRun frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
Colibrì is a pure-C, zero-dependency inference engine that runs frontier MoE models (GLM-5.x 744B, Kimi K3 2.8T, DeepSeek V4, Qwen3.x) on consumer hardware by streaming disk-resident experts into a unified VRAM/RAM/NVMe tier hierarchy. It offers chat/serve/web front ends plus CPU, CUDA, Metal and Vulkan backends.
Use Cases
Run 744B-2.8T MoE models on consumer hardwareLocal private LLM inference without API providersStreaming routed experts from NVMe/SSD on demandMultitier VRAM/RAM/disk weight placementOpenAI-style local chat and serve endpoints (chat/serve/web)CPU-only inference with no GPUMulti-GPU and heterogeneous CPU/CUDA/Metal executionDisk-backed local cluster inference across machinesMoE expert routing research and benchmarkingQuantized int4/fp8 weight conversion pipelinesLong-context inference with compressed MLA KV stateSpeculative decoding with MTP and grammar drafts
Built With
- Language
- C
- Frameworks
- CUDA · Metal · Vulkan · PyTorch · Hugging Face Hub · NumPy · Transformers · safetensors
Tags
inference-engine · moe · local-llm · c · zero-dependencies · quantization · int4 · expert-streaming · cpu-inference · gpu-backends · kv-cache · speculative-decoding · disk-offloading · llm-serving · nuive-streaming · memory-hierarchy