Vibe Coding Discover

AI Frameworks

ninfer-fusion-kvmem

View on GitHub
★ 26221 forksC++Apache-2.01314521gjy

A Windows/NVIDIA-focused NInfer inference-engine fork that uses host-backed KV storage and content-based retrieval to support logical contexts larger than the device KV pool. Includes serving, build, tuning, and verification tools; the experimental implementation has documented reliability and quality limitations.

Use Cases

Run long-context LLM inference on NVIDIA GPUsReduce device KV-cache memory use with host-backed storageServe local models through an OpenAI-compatible APIReuse cached context across turnsExperiment with speculative decoding

Built With

Language
C++
Frameworks
NInfer · CUDA

Tags

LLM inference · long context · KV cache · GPU memory · NVIDIA · Windows · CUDA · host-backed storage · speculative decoding · OpenAI-compatible API