LMCache
View on GitHubLMCache: Supercharge Your LLM with the Fastest KV Cache Layer
LMCache is a vendor-neutral KV cache management layer for LLM inference. It stores and reuses KV cache across CPU RAM, disk and remote backends, cutting TTFT and boosting throughput for long-context, multi-turn agentic and RAG workloads on engines like vLLM.
Use Cases
Reduce time-to-first-token for long-context LLM servingReuse KV cache across requests, sessions and engine instancesOffload KV cache to CPU RAM, SSD, Redis or object storageSpeed up multi-turn agentic and chat workloadsAccelerate RAG / knowledge-augmented pipelines via CacheBlend non-prefix reusePrefill-decode disaggregation with KV transfer over NVLink/RDMA/TCPMonitor KV cache hit rates and lifecycle metrics in productionImprove MoE inference throughputCut repeated prefill compute cost for repeated promptsAdd a standalone KV cache daemon to existing inference engines
Built With
- Language
- Python
- Frameworks
- vLLM · PyTorch · NIXL · Redis · Valkey · Mooncake · InfiniStore · S3 · gRPC · Kubernetes · CUDA · ROCm · SGLang · TensorRT-LLM
Tags
kv-cache · llm-inference · inference-optimization · vllm · prefix-caching · cache-offloading · ttft · throughput · long-context · observability · pd-disaggregation · tiered-storage · cuda · rocm · vendor-neutral · serving-engine