LMCache

KV cache layer that reuses LLM caches across GPU, CPU, disk, and object storage to cut TTFT.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

Recomputing a long prompt's KV cache on every request is the tax LMCache removes. The library adds a cache management layer to LLM serving engines that persists key-value tensors across GPU memory, CPU RAM, local disk, and remote backends including Redis, Valkey, and S3-compatible object storage, then reloads them wherever the next request lands, cutting time-to-first-token sharply for long-context, multi-turn, RAG, and agentic workloads; recent releases report gains up to 10x on MoE-heavy inference. Integration is first class with vLLM and the vLLM production stack, SGLang, and NVIDIA's Dynamo datacenter framework, and hardware support spans NVIDIA, AMD MI300X, Arm, and Ascend. A pip install of lmcache plus a few lines of configuration attaches it to an existing deployment, with documentation maintained at docs.lmcache.ai. Governance moved under the PyTorch Foundation in October 2025, and the Apache 2.0 codebase has drawn over 11,000 GitHub stars, making it a default answer for teams whose GPU bills are dominated by redundant prefill.

Should you use LMCache?

Pick it when

Pick LMCache when your vLLM or SGLang deployment serves long prompts, multi-turn chat, RAG, or agents that resend the same context and prefill dominates GPU time, so cached KV tensors in RAM, disk, Redis, or S3 can be reused.

Look elsewhere when

Skip it for short, unique prompts where little context repeats, or for single-user setups on llama.cpp or Ollama, which it does not target. If you need routing and autoscaling as well, llm-d or NVIDIA Dynamo include cache tiers.

Alternatives to LMCache

  • llm-d

    A full Kubernetes stack with cache-aware routing and tiered offloading, more than a cache layer but far more to deploy and operate.

  • NVIDIA Dynamo

    Adds KV-aware routing, disaggregation, and its own tiered offload across many nodes; heavier to adopt, and it can use LMCache underneath.

  • SGLang

    Built-in RadixAttention reuses prefixes in one server's GPU memory, often enough before adding LMCache's cross-node, multi-tier storage.

  • AIBrix

    Includes a distributed KV cache inside a broader vLLM control plane on Kubernetes, if you want routing and autoscaling in one package.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Added
Aug 24, 2026

Tags