1. Set the service target
Enter the aggregate token throughput the deployment must sustain.
2. Use measured GPU throughput
Provide tokens per second achieved by one GPU under a representative model, precision, batch, and context mix.
3. Choose sustainable utilization
Reduce benchmark throughput to a level that leaves operating margin.
4. Add redundancy
Enter spare devices required for maintenance or failure tolerance.
5. Check memory fit
Provide usable memory per GPU and the combined model-plus-runtime memory footprint.
6. Review the binding constraint
Compare throughput-driven and memory-driven counts; the larger requirement determines the base cluster size.