Model Overview

The stack contains every transformer layer. Prefill processes the prompt tokens together. Decode processes one new token at a time. Start the animation, then select a layer to inspect its operations.

Model

Hardware

Quantization

Request

Inference Phase

Summary

Data Movement During Inference

Transformer Layer Stack

Model Totals (per decode token, batch=1)

Click a preset to change the settings.