KV Cache & Memory

During generation, attention reuses key and value data from earlier tokens. The KV cache stores this data. Its size grows with the sequence length and can limit the batch size.

Model

Hardware

Request

Attention Variant

Precision

Summary

Weights
KV Cache
Total
GPU Memory

GPU Memory Budget

KV Cache Formula

Attention Variant Comparison

Does It Fit?

GPUVRAMWeightsKV CacheTotalFits?GPUs Needed

KV Cache vs Sequence Length

Click a preset to change the settings.