1. Enter peak concurrency
Use concurrent sessions after traffic and latency capacity planning.
2. Add benchmark density
Enter the number of sessions one GPU sustained at the required quality and latency.
3. Set utilization headroom
Choose a target below full saturation to protect tail latency.
4. Add redundancy
Reserve capacity for failures, maintenance, or zone loss.
5. Include stack overhead
Add capacity for embedding, speech, routing, or other GPU workloads not included in the benchmark.
6. Review rounded count
Deploy whole GPUs and validate the result under realistic load.