1. Measure peak arrivals
Enter the highest sustained request rate the service should handle, not a brief one-second spike unless that spike must be served without queueing.
2. Enter observed latency
Use end-to-end time from request acceptance to completed response for the target model and payload mix.
3. Set utilization
Choose a worker utilization target below 100% to preserve scheduling and recovery margin.
4. Add traffic headroom
Increase the arrival rate for forecast error, bursts, or growth.
5. Enter current workers
Provide the number of concurrent execution slots already available.
6. Compare capacity and objective
Review required workers, estimated request capacity, and whether average latency meets the stated objective.