1. Step 1
Enter the peak arrival rate in tasks per minute for the service boundary being sized.
2. Step 2
Use measured end-to-end average latency for one task, including model and required orchestration time.
3. Step 3
Add headroom for traffic variation, retries, and imperfect load balancing.
4. Step 4
Enter currently available concurrent slots to compare the plan with existing capacity.
5. Step 5
Review the rounded required slots and the surplus or shortfall.
6. Step 6
Run a separate scenario for a high-percentile latency when stricter service protection is needed.