Layer Overview

One transformer layer contains attention, a feed-forward network (FFN), and supporting operations. The diagram follows token data through each operation. Select a block to inspect its matrix size, arithmetic work, and memory traffic.

Model

Hardware

Request

Performance Options

Selected Operation

Name
Category
Matrix (M×K×N)
FLOPs
Traffic
FLOPs/Byte
Limit
Time
Inspect tiling for this operation →

Transformer Layer Operation Flow

FLOPs Breakdown (per layer)

Memory Traffic Breakdown (per layer)

All Operations

Operation Type M K N FLOPs Traffic FLOPs/Byte Limit Time (ms) % FLOPs

Roofline Model

Click a preset to change the settings.