1. Set the completion target
Enter the sustained successful image tasks required each hour.
2. Use measured GPU time
Benchmark the selected model, resolution, batch size, and precision mode.
3. Enter parallel jobs
Specify how many attempts one GPU can process concurrently without changing the measured latency materially.
4. Add retry overhead
Include failed or repeated generations that consume GPU time.
5. Apply utilization and reserve
Choose a sustainable utilization target and redundancy allowance for maintenance or failures.