ninfer-fusion-kvmem
View on GitHubA Windows/NVIDIA-focused NInfer inference-engine fork that uses host-backed KV storage and content-based retrieval to support logical contexts larger than the device KV pool. Includes serving, build, tuning, and verification tools; the experimental implementation has documented reliability and quality limitations.
Use Cases
Run long-context LLM inference on NVIDIA GPUsReduce device KV-cache memory use with host-backed storageServe local models through an OpenAI-compatible APIReuse cached context across turnsExperiment with speculative decoding
Built With
- Language
- C++
- Frameworks
- NInfer · CUDA
Tags
LLM inference · long context · KV cache · GPU memory · NVIDIA · Windows · CUDA · host-backed storage · speculative decoding · OpenAI-compatible API