LLM Inference Explorer
Model mechanics
Model
Layer
Tokens
Embed
Position
Attention
Patterns
Decoding
Generation
Shapes
Performance
Powers of Ten
GEMM
Flash Attention
KV Cache
Batching
Speculative Decode
Cost
Decoding
The model outputs
logits
for every token in its vocabulary.
Softmax
converts these to probabilities, then a
sampling strategy
selects the next token.
Temperature
T = 1.0
0.1 (peaked)
1.0 (default)
3.0 (flat)
Sampling Strategy
Greedy
Top-k
Top-p
Top-k + p
Top-k
5
Top-p
0.90
Statistics
Top-1 Prob
—
Candidates
—
Filtered
—
Entropy
—
bits
Eff. Vocab
—
tokens
Strategy
—
Top Tokens
Logits (fixed raw model scores)
Sampling Distribution
Samples
Resample
Presets
Explore
Questions
What Next
Tips
▼
Click a preset to change the settings.
Greedy
Nucleus (p=0.9)
Top-k (k=5)
High temperature
Set top-k to 5 and click Resample repeatedly. How many distinct tokens appear?
Compare top-p=0.9 at T=0.5 and T=2.0. How does the candidate count change?
Why does greedy always produce the same token, while top-p gives variety?
Set top-k=20 (all tokens). How does it compare to no filtering?
Why does top-p adapt to the distribution shape while top-k does not?
If T=0.1 makes the distribution very peaked, does top-k even matter at low temperature?
How do top-k and top-p interact when combined? Which filter is applied first?
Autoregressive Generation
repeats next-token selection one token at a time.
Attention
uses softmax to compute weights rather than token probabilities.