LMCache
KV cache layer that reuses LLM caches across GPU, CPU, disk, and object storage to cut TTFT.
About
Recomputing a long prompt's KV cache on every request is the tax LMCache removes. The library adds a cache management layer to LLM serving engines that persists key-value tensors across GPU memory, CPU RAM, local disk, and remote backends including Redis, Valkey, and S3-compatible object storage, then reloads them wherever the next request lands, cutting time-to-first-token sharply for long-context, multi-turn, RAG, and agentic workloads; recent releases report gains up to 10x on MoE-heavy inference. Integration is first class with vLLM and the vLLM production stack, SGLang, and NVIDIA's Dynamo datacenter framework, and hardware support spans NVIDIA, AMD MI300X, Arm, and Ascend. A pip install of lmcache plus a few lines of configuration attaches it to an existing deployment, with documentation maintained at docs.lmcache.ai. Governance moved under the PyTorch Foundation in October 2025, and the Apache 2.0 codebase has drawn over 11,000 GitHub stars, making it a default answer for teams whose GPU bills are dominated by redundant prefill.
Should you use LMCache?
Pick it when
Pick LMCache when your vLLM or SGLang deployment serves long prompts, multi-turn chat, RAG, or agents that resend the same context and prefill dominates GPU time, so cached KV tensors in RAM, disk, Redis, or S3 can be reused.
Look elsewhere when
Skip it for short, unique prompts where little context repeats, or for single-user setups on llama.cpp or Ollama, which it does not target. If you need routing and autoscaling as well, llm-d or NVIDIA Dynamo include cache tiers.
Alternatives to LMCache
- llm-d
A full Kubernetes stack with cache-aware routing and tiered offloading, more than a cache layer but far more to deploy and operate.
- NVIDIA Dynamo
Adds KV-aware routing, disaggregation, and its own tiered offload across many nodes; heavier to adopt, and it can use LMCache underneath.
- SGLang
Built-in RadixAttention reuses prefixes in one server's GPU memory, often enough before adding LMCache's cross-node, multi-tier storage.
- AIBrix
Includes a distributed KV cache inside a broader vLLM control plane on Kubernetes, if you want routing and autoscaling in one package.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- LLM Inference & Serving
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026