1. Enter peak workload
Use the busiest sustained hour rather than a daily average.
2. Enter token demand per task
Include all input and output tokens processed by the serving fleet.
3. Add measured GPU throughput
Use an effective benchmark for the chosen model, hardware, precision, sequence length, and batching setup.
4. Set utilization
Choose the share of measured throughput you are willing to plan against.
5. Add resilience
Enter extra capacity for bursts, failures, rolling updates, and forecast uncertainty.
6. Review provisioned GPUs
The result rounds up because partial GPUs cannot normally be provisioned independently.