Case study

Where the KV cache lives: storage tiering for LLM inference under real-world traffic

LLM serving runs out of GPU memory long before it runs out of compute. Using real conversation and production request traces, I measured what each storage tier below the GPU costs a live request and built a shared cache tier on GekkoFS that lets several vLLM instances reuse each other's work.

Context
Research project, JGU Mainz
Role
Researcher and engineer, end to end
Domain
LLM inference, HPC storage, MLOps
Infrastructure
Mogon HPC cluster (A40, A100), workstation with 2x RTX A6000
Models
Llama 3.1 8B and 70B, Qwen3.5-35B-A3B
Year
2026
Results
Latency
Reloading beats recomputing by 3.5x to 4.6x

Reading an evicted prefix back from CPU memory took 73 ms on an A100, where recomputing it took 256 ms. The slower the GPU, the larger the saving, so cache tiering pays off most on cost-efficient hardware.

Architecture
Latency, not bandwidth, sets the cost

Controlled sweeps separated two properties that earlier evaluations change together. Serving survived a cut to a quarter of the disk bandwidth, but one millisecond of added per-operation latency multiplied time to first token several times over. Decode latency never moved.

Scale-out
Shared cache at parity with private NVMe

With four vLLM instances on two nodes sharing one GekkoFS namespace, an instance reloads KV that another instance computed instead of recomputing it. After tuning, the shared tier tracks the private local-disk baseline.

Tuning
21x faster warm reload on the distributed tier

Tuning the backend and raising the chunk size cut the reload of a 16k-token document to about 1.9 s and let every workload cell finish. With the default setup the largest cells could not complete.

The story

Every token an LLM generates depends on key and value vectors for all tokens before it. vLLM keeps these in a KV cache so it does not recompute them, and that cache grows with context length and concurrency until GPU memory is full. From there it spills to CPU DRAM and then to NVMe, and the storage path starts deciding how fast the first token comes back. Published systems show that offloading works. Few say what the storage underneath has to deliver.

I built the experiment platform on vLLM 0.19 with LMCache as the tiering layer, containerized with Apptainer on the Mogon cluster and Docker on a GPU workstation, and drove it with ShareGPT and BurstGPT traffic. Fourteen experiments moved from a single GPU up to production-style traffic on a 70B model with tensor parallelism. I isolated the storage variable with a FUSE throttling shim that caps bandwidth and injects per-operation latency independently, and cross-checked the ordering against a kernel-level dm-delay device.

Stack of four storage tiers from GPU HBM down to a shared GekkoFS tier, with the prefill read path in bold and the decode path bypassing storage.
The KV cache spills down the hierarchy in steps. The cost lives on the read path during prefill, and decode never touches storage.
Experiment pipeline from workload generator through vLLM and LMCache to a throttled disk tier, with separate bandwidth and latency controls and a metrics path.
The sensitivity testbed varies bandwidth and per-operation latency independently, which whole-tier comparisons cannot do.

What it means

The results locate the cost precisely. Writing the cache out is close to free in steady state and decode never touches storage, so the whole penalty lands on prefill. Within prefill, per-operation latency binds and bandwidth barely matters. That result reads as a requirements specification for any tier placed under the cache: a fraction of one NVMe device's bandwidth is enough, but every millisecond of round-trip time is paid in full.

Networked storage is exactly where latency is hard to keep low, so I integrated GekkoFS, an ad-hoc file system built at job start from the allocation's own node-local NVMe. I wrote SharedFSBackend, which plugs it into LMCache as a shared disk tier. On a single instance it lost to local NVMe, reading at about 0.7 GB/s against 6 GB/s with nothing to hide the round trip behind. Its value appeared once several instances served behind a load balancer. A private cache means each replica recomputes work its neighbors already did, whereas one shared namespace stores each prefix once and lets any replica read it.

Two nodes with two vLLM instances each, a load balancer in front, and one GekkoFS namespace spanning the node-local NVMe of both nodes.
Four instances share one cache namespace, so a prefix computed on one node is reloaded on the other instead of recomputed.

Design rule

For an inference platform the design rule is concrete. Keep the hot tiers node-local and judge any shared tier by its per-operation latency before its throughput. A shared tier earns its place once the fleet runs more than one replica and users share prefixes. The 8B model used for most runs is the hardest case for a reload to win, so the parity measured here is a floor. At 70B, where recomputation costs far more per token, the same design should turn parity into a latency win.

AWS reference architecture with GPU nodes on EKS in a cluster placement group, instance-store NVMe, FSx for Lustre as shared tier, a prefix-aware router, and observability.
Reference architecture: how the measured design rules map onto AWS services.

Technology

Inference
vLLM
vLLM
Serving engine with PagedAttention and tensor parallelism
LM
LMCache
KV cache tiering across CPU memory, disk and shared backends
Hugging Face Hub
Hugging Face Hub
Model hosting for Llama 3.1 and Qwen3.5
Models
Llama 3.1 8B Instruct
Llama 3.1 8B Instruct
Single-GPU baseline, 128 KiB of KV cache per token
Llama 3.1 70B Instruct
Llama 3.1 70B Instruct
Multi-GPU case with tensor parallelism, 320 KiB of KV cache per token
Qwen3.5-35B-A3B (FP8)
Qwen3.5-35B-A3B (FP8)
Long-context case, hybrid attention with 20 KiB per token
Storage
GK
GekkoFS
Ad-hoc distributed file system as shared cache tier
</>
SharedFSBackend
My LMCache backend for shared file system tiers
FS
FUSE throttling shim
Independent bandwidth caps and latency injection
dm-delay
dm-delay
Kernel-level latency injection for cross-checks
Compute and operations
Slurm
Slurm
Job scheduling on the Mogon cluster
AP
Apptainer
Reproducible containers on HPC nodes
Docker
Docker
Containers on the local GPU workstation
NVIDIA A100 and A40
NVIDIA A100 and A40
GPU testbeds with different compute-to-memory ratios
AWS reference mapping
EC2 P4d / P5
EC2 P4d / P5
GPU nodes with local NVMe instance store
FSx for Lustre
FSx for Lustre
Managed shared cache tier across replicas
EFA and cluster placement groups
EFA and cluster placement groups
Low round-trip latency to the shared tier
EKS with Karpenter
EKS with Karpenter
EKS with Karpenter
GPU replica scaling behind a prefix-aware router
AWS ParallelCluster
AWS ParallelCluster
Slurm-based alternative that mirrors the HPC setup
S3 and ECR
S3 and ECR
S3 and ECR
Model weights and container images
Managed Prometheus and Grafana
Managed Prometheus and Grafana
Managed Prometheus and Grafana
TTFT and inter-token latency SLOs
Core stack

Working on something similar?

Tell me what you run today and where it hurts. I will come back with how I would build it, and what I would leave alone.

Discuss your project