1. Enter daily document demand
Use the expected completed volume, not only peak arrivals.
2. Count model tasks per document
Include classification, extraction, validation, or other separate calls.
3. Measure average task latency
Use end-to-end latency from representative requests.
4. Choose target utilization
A lower percentage provides more queue and failure headroom.
5. Set the processing window
Enter the hours available to clear the day’s workload.
6. Review required workers
Round up is applied because partial workers cannot supply full concurrency.