Batching Simulator

During decode, a batch of one loads all model weights for one token. This uses little of the GPU's arithmetic capacity. A larger batch shares each weight load across more tokens. The FLOPs-per-byte ratio, called arithmetic intensity, increases.

Model

Hardware

Request

Quantization

GPU Arithmetic Use

Throughput & Latency

Arithmetic Intensity vs Batch Size

Static vs Continuous Batching

Memory Budget

Click a preset to change the settings.