Generation

Autoregressive generation produces one token at a time. Each token depends on the tokens before it. Prefill processes all prompt tokens together. Decode then produces one new token per pass through the model.

Model

Hardware

Prompt

Controls

Phase

Idle

Token Sequence

KV Cache

Timing

Click a preset to change the settings.