1. Enter active GPUs
Count GPUs available to the translation service during the planning window.
2. Provide measured token speed
Use a benchmark from the same model, precision, and request profile.
3. Estimate tokens per attempt
Include source, generated output, and recurring prompt context.
4. Apply efficiency and utilization
Reflect batching behavior and sustainable operating load.
5. Include retries and backlog
Account for repeated calls and optionally estimate how long a queue will take to clear.