ninfer-fusion-kvmem
★ 262A Windows/NVIDIA-focused NInfer inference-engine fork that uses host-backed KV storage and content-based retrieval to support logical contexts larger than the device KV pool. Includes serving, build, tuning, and verification tools; the experimental implementation has documented reliability and quality limitations.
AI Frameworks | C++ · LLM inference · long context
View Project →