The stack contains every transformer layer. Prefill processes the prompt tokens together. Decode processes one new token at a time. Start the animation, then select a layer to inspect its operations.
Model
Hardware
Quantization
Request
Inference Phase
Summary
Data Movement During Inference
Transformer Layer Stack
Model Totals (per decode token, batch=1)
Click a preset to change the settings.
Load GPT-2, which has only 12 layers. Now switch to 405B to compare their scale.
Select any transformer layer to inspect its operations.