1. Enter GPU count
Provide the number of devices actively serving the model.
2. Add per-GPU performance
Use a measured token rate from the intended model, precision, and serving engine.
3. Apply scaling efficiency
Account for synchronization, communication, scheduler, and load-balancing losses.
4. Describe the token mix
Set the fraction of total tokens attributable to input processing.
5. Enter request size
Use average input plus generated tokens per request.
6. Review sustained output
The daily figure applies the selected availability percentage to a full 24-hour period.