Presets
Training a 7.08B dense transformer ▼ model in BF16 ▼ on 1,024 × H100 SXM ▼ with 10²⁴ FLOPs takes 28.6 days and uses 21.8T tokens.
predicted loss 1.95 · between Llama 3 8B and Chinchilladashed = guessed sizes
01
C ≈ 6 · N · D
6 × 7.08B active params × 21.8T tokens (attention adds 7% at 4.1K context)
10²⁴ FLOPs
02
T = C ÷ (n · F · MFU)
10²⁴ ÷ (1,024 × 989 TFLOP/s × 40%)
28.6 days
03
F ÷ BW
989 TFLOP/s ÷ 3.35 TB/s
295 FLOP/byte
04
2BDF ÷ bytes(BD + DF + BF) ≈ B
batch 1 → 1 FLOP/byte, ridge is 295
memory-bound
05
N_max = n · M ÷ 16
1,024 × 80 GB ÷ 16 bytes per param
5.12T params
06
tok/s ≈ BW ÷ (bytes · N_active)
8 × 3.35 TB/s ÷ 14.4 GB per token
1,855 tok/s
Hardware cost$30.7M1,024 × $30K
Energy590 MWh≈ 56.2 US homes for a year
Electricity bill$59Kat $0.10 per kWh
1,855 tok/sbatch 1 · 8-way parallel
320B paramsat BF16 on 8 devices