1. Set traffic demand
Enter the sustained requests per second to be served.
2. Describe cache performance
Provide the hit rate plus GPU processing time for hits and misses.
3. Reserve operating headroom
Set the fraction of theoretical GPU capacity considered safely usable.
4. Use the rounded result
Provision at least the displayed whole-GPU count, then verify with a load test.