LLM Inference Explorer
Model mechanics
Model
Layer
Tokens
Embed
Position
Attention
Patterns
Decoding
Generation
Shapes
Performance
Powers of Ten
GEMM
Flash Attention
KV Cache
Batching
Speculative Decode
Cost
Speculative Decoding
A small
draft model
proposes K tokens. The
target model
checks all K in one pass. Accepted draft tokens reduce the number of target-model steps.
Target Model
Draft Model
Hardware
Parameters
Draft tokens (K)
4
Acceptance rate (α)
0.80
Speculative Decoding Cycle
Ready
Draft
Verify ↓
Target
Play
Step
Reset
Speed
5
Output Sequence
Tokens will appear here as cycles complete.
—
expected speedup
Time Breakdown per Cycle
Speedup vs Acceptance Rate
Presets
Explore
Questions
What Next
Tips
▼
Click a preset to change the settings.
High α
α=0.80
Low α
K=2, low acceptance
Run at K=4, α=0.80
and step through the animation. Watch the accept/reject pattern.
Drag α from 0.95 down to 0.30. At what point does the speedup drop below 1.1×?
Set K=6, α=0.40
, then
try K=2, α=0.40
. The shorter draft is faster at low acceptance.
Try different draft models. How does the draft-to-target size ratio affect speedup?
Why can the target model verify K draft tokens in the same time as generating 1 token?
At what acceptance rate does K=2 become faster than K=6? Why is the shorter draft faster when acceptance is low?
Cost Estimator
estimates inference time and cost.