1. Enter target completions
Use the sustained number of successfully completed image requests per minute.
2. Measure end-to-end processing latency
Enter average active processing time, excluding time a request waits in an external queue.
3. Include extra attempts
Add retries or regenerated attempts that consume the same worker pool.
4. Set available slots
Count simultaneous tasks the current deployment can process.
5. Choose a utilization ceiling
Keep this below 100% to preserve room for variability and bursts.