1. Enter monthly run volume
Count successful and failed runs if both consume resources.
2. Estimate average runtime
Use elapsed compute time per run, not human waiting time.
3. Set the compute rate
Enter the effective hourly infrastructure or execution charge.
4. Add token usage
Use total input and output tokens consumed by an average run.
5. Enter the token price
Use a blended cost per one million tokens when multiple models are involved.
6. Add tool charges
Include search, browser, database, messaging, or other per-run API costs.
7. Review the monthly total
Inspect each component to identify the largest cost driver.