kimi-k3-in-c
★ 8.3KA portable C99 inference engine that runs the 2.78T-parameter Kimi K3 MoE model on one CPU with 8.24 GB peak RAM, streaming a 1.56 TB checkpoint from disk with no BLAS, framework, GPU or runtime dependencies.
AI Frameworks | C · llm-inference · cpu-inference
View Project →