Vibe Coding Discover

Use Cases

Memory-budget Tuning Of Model Streaming And Expert Lru Caching

Published projects tagged with this use case.

1 project

kimi-k3-in-c

★ 8.3K

A portable C99 inference engine that runs the 2.78T-parameter Kimi K3 MoE model on one CPU with 8.24 GB peak RAM, streaming a 1.56 TB checkpoint from disk with no BLAS, framework, GPU or runtime dependencies.

AI Frameworks | C · llm-inference · cpu-inference

View Project →