1. Enter the training workload
Provide tokens per epoch and the planned number of epochs.
2. Use per-GPU throughput
Enter a benchmark for the intended model, precision, sequence length, and hardware.
3. Apply scaling efficiency
Reduce the ideal linear gain to reflect communication and synchronization overhead.
4. Set utilization and deadline
Reserve operational headroom and enter the desired completion time.
5. Provision the rounded count
Use at least the displayed whole-GPU count, subject to memory and topology constraints.