Layer Overview

One transformer layer contains attention, a feed-forward network (FFN), and supporting operations. The diagram follows token data through each operation. Select a block to inspect its matrix size, arithmetic work, and memory traffic.

Model

Hardware

Request

Performance Options

Selected Operation

Name
—
Category
—
Matrix (M×K×N)
—
FLOPs
—
Traffic
—
FLOPs/Byte
—
Limit
—
Time
—
Inspect tiling for this operation →

Transformer Layer Operation Flow

FLOPs Breakdown (per layer)

Memory Traffic Breakdown (per layer)

All Operations

Operation Type M K N FLOPs Traffic FLOPs/Byte Limit Time (ms) % FLOPs

Roofline Model

Click a preset to change the settings.