How much model can a compute budget buy?

Back-of-envelope math from Stanford’s CS336.

Presets

Training a 7.08B dense transformer ▼ model in BF16 ▼ on 1,024 × H100 SXM ▼ with 10²⁴ FLOPs takes 28.6 days and uses 21.8T tokens.
1,024 × H100 SXMn · F = 1.01 EFLOP/sL = 33d_model = 4,227N = 7.08BD = 21.8T tokens≈ 168M booksHARDWAREMODELTRAINING DATAChinchilla: 20 × N

predicted loss 1.95 · between Llama 3 8B and Chinchilla
10²⁰10²¹10²²10²³10²⁴10²⁵10²⁶10²⁷1.751.801.902.002.252.503.003.50compute-optimal frontierGPT-2TinyLlamaPhi-3 miniGPT-3ChinchillaLlama 3 8BDeepSeek V3Llama 3 70BGPT-4Llama 3.1 405BGPT-5Claude Fable 5GPT-6 AstraSelected modeltraining compute (FLOPs) →↑ lower loss
dashed = guessed sizes

01
C ≈ 6 · N · D
6 × 7.08B active params × 21.8T tokens (attention adds 7% at 4.1K context)
10²⁴ FLOPs
02
T = C ÷ (n · F · MFU)
10²⁴ ÷ (1,024 × 989 TFLOP/s × 40%)
28.6 days
03
F ÷ BW
989 TFLOP/s ÷ 3.35 TB/s
295 FLOP/byte
04
2BDF ÷ bytes(BD + DF + BF) ≈ B
batch 1 → 1 FLOP/byte, ridge is 295
memory-bound
05
N_max = n · M ÷ 16
1,024 × 80 GB ÷ 16 bytes per param
5.12T params
06
tok/s ≈ BW ÷ (bytes · N_active)
8 × 3.35 TB/s ÷ 14.4 GB per token
1,855 tok/s

Hardware cost$30.7M1,024 × $30K
Energy590 MWh≈ 56.2 US homes for a year
Electricity bill$59Kat $0.10 per kWh
1,855 tok/sbatch 1 · 8-way parallel
320B paramsat BF16 on 8 devices