1. Enter fleet size
Provide the number of GPU devices available to the model-serving workload.
2. Use representative throughput
Enter measured token throughput per GPU for the actual model and serving configuration.
3. Reserve operating headroom
Set the target utilization below 100% to leave room for variation and operational work.
4. Enter task size
Use combined input and output tokens per completed task.
5. Add peak demand
Enter expected tasks per hour to compare capacity with workload.
6. Review capacity
Check both tasks per hour and demand coverage; coverage below 100% indicates a likely queue buildup.