Turn a training plan into compute, time and energy.
Define the cluster, useful compute utilization, average power draw and facility efficiency, then compare them with an approximate dense-pretraining workload. Every assumption that materially changes the result is editable.
Keep delivered compute separate from energy consumed.
—
From IT power to facility energy
The bars separate accelerators, other IT and facility overhead introduced by PUE. TDP is not treated as measured draw: average power fraction is controlled above.
Power and energy
Average facility power: —
Energy during scheduled window: —
Energy to complete approximate workload: —
Effective throughput
Sustained model FLOPs: —
Model FLOPs per facility MW: —
MFU and average power draw are different variables: a cluster can consume substantial power while converting only a modest fraction of theoretical peak into useful model FLOPs.
Three layers, kept separate.
The workload uses C≈kND with k=6 by default. It is a dense-pretraining approximation, not an exact measurement for every architecture.
Useful throughput is accelerators × dense peak × MFU. MFU stays editable because it depends on architecture, parallelism, sequence length, software and cluster scale.
Accelerator power uses TDP × average power fraction. Other IT is added next, then PUE converts IT energy into total facility energy.
This is an engineering estimator, not an electricity bill or a carbon measurement.
TDP is a configurable design limit, not a power reading. MFU is not electrical utilization. PUE is a facility metric and can vary with load and climate. The 6ND approximation omits attention, vocabulary, MoE, recomputation and other training details. Energy is not converted into emissions because that requires reliable time- and location-specific grid carbon intensity.