Attention Patterns

An attention mask marks which earlier token positions each token can use. The grid shows the allowed token pairs. A sparse GPU operation can skip fully masked blocks. The KV cache stores keys and values from earlier tokens.

Pattern

Sequence

Pattern Settings

KV Cache

Animation

Query × Key Mask Grid

computed masked partial block current query

Dynamic KV Cache

Cost Signals

Mask Rule

Click a preset to change the settings.