1. Enter hourly task demand
Use successful translation tasks expected during the peak sustained hour.
2. Estimate tokens per attempt
Include source, output, and repeated instruction tokens processed by the model.
3. Use a production-like benchmark
Enter measured tokens per second for one GPU under the target model and precision.
4. Adjust batching and utilization
Apply observed batching efficiency and a sustainable utilization target.
5. Add retries and reserve
Include repeated model calls and spare capacity for failures or maintenance.