Vibe Coding Discover

AI Frameworks

kimi-k3-in-c

View on GitHub

A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

★ 8.3K1,331 forksCApache-2.0FareedKhan-dev

A portable C99 inference engine that runs the 2.78T-parameter Kimi K3 MoE model on one CPU with 8.24 GB peak RAM, streaming a 1.56 TB checkpoint from disk with no BLAS, framework, GPU or runtime dependencies.

Use Cases

Run a 2.78T-parameter MoE LLM on a single CPU with 8 GB RAMCPU-only local LLM inference without GPU or BLASStudy a from-scratch transformer/MoE inference engine in portable C99Memory-budget tuning of model streaming and expert LRU cachingBenchmark disk-streaming vs resident inference performanceVerify C kernels against a PyTorch reference via fixtures

Built With

Language
C
Frameworks
OpenMP · CMake · Make · PyTorch · NumPy · tiktoken

Tags

llm-inference · cpu-inference · c99 · zero-dependencies · quantization · mxfp4 · mixture-of-experts · moe · avx2 · simd · memory-efficient · linear-attention · kv-cache · streaming-weights · from-scratch · safetensors

kimi-k3-in-c — Vibe Coding Discover