1. Enter available concurrency
Use the number of requests that can be processed at the same time under your service or deployment limits.
2. Measure end-to-end latency
Include model processing plus any synchronous retrieval, tool, moderation, or validation steps.
3. Set sustainable utilization
Leave headroom for variability rather than using 100% unless the workload is fully controlled.
4. Add a target rate
Enter the task arrival or generation rate the system should sustain.
5. Define a workload size
Use the planned batch size or the number of AI turns expected during a peak hour.
6. Compare capacity and demand
A negative gap indicates that demand exceeds the estimated sustainable rate.