1. Enter a per-GPU benchmark
Use a measurement that matches the model, precision, batch size, and sequence length.
2. Set the cluster size
Enter the number of GPUs assigned to the run.
3. Apply scaling efficiency
Account for communication and synchronization losses.
4. Apply utilization
Reflect stalls, data loading, checkpointing, and other non-productive time.
5. Add total token workload
The calculator converts effective throughput into estimated run duration.