1. Enter workflow traffic
Use expected completed multi-agent workflows per minute.
2. Describe agent activity
Enter model-backed agents, calls per agent, and generated tokens per call.
3. Estimate the parallel peak
Choose the share of calls likely to overlap across agents.
4. Use measured GPU throughput
Enter effective output throughput for one GPU at the intended configuration.
5. Add headroom
Set utilization and reserve, then review average and peak-adjusted demand.